科技回声

4 条评论

> Previous State-of-the-Art: [...] The number of parameters in the LSTM layers of these models vary from 2 million to 151 million.> We present model architectures in which a MoE with up to 137 billion parametersBack in 2017 most models were well under 1B, GPT2 (2019) was one of the first "big" non-MOE models at 1.5B in size. People weren't sure how well/much they would scale.The CoralAI TPU has a mere 8 MB of SRAM in 2019!GPT3 was 175B in 2020.Now nearly all LLM's are at minimum 1B, but dense 70B is now common.

评论 #38574779 未加载

dang超过 1 年前

Discussed at the time:Outrageously Large Neural Networks: The Sparsely-Gated Mixture-Of-Experts Layer - <a href="https://news.ycombinator.com/item?id=13518039">https://news.ycombinator.com/item?id=13518039</a> - Jan 2017 (81 comments)Outrageously large neural networks: the sparsely-gated mixture-of-experts layer - <a href="https://news.ycombinator.com/item?id=12963364">https://news.ycombinator.com/item?id=12963364</a> - Nov 2016 (2 comments)

gryfft超过 1 年前

[2017]

评论 #38573015 未加载

评论 #38573502 未加载

legel超过 1 年前

In response to a very impressive set of LLM open weights released today from Mistral AI called "Mixtral 8x7B", I was reminded of this amazing publication on the origin of "sparse mixture of experts" from none other than Geoffrey Hinton and Jeff Dean.The Sparse Mixture of Experts neural network architecture is actually an absolutely brilliant move here. It scales fantastically, when you consider that (1) GPU RAM is way too expensive, in financial dollars, (2) SSD / CPU RAM are relatively cheap, and (3) you can have "experts" running on their own computers, i.e. it's a natural distributed computing partitioning strategy for neural networks.I did my M.S. thesis on large-scale distributed deep neural networks in 2013 and can say that I'm delighted to point out where this came from.In 2017, it emerged from a Geoffrey Hinton / Jeff Dean / Quoc Le publication called "Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer".Here is the abstract: "The capacity of a neural network to absorb information is limited by its number of parameters. Conditional computation, where parts of the network are active on a per-example basis, has been proposed in theory as a way of dramatically increasing model capacity without a proportional increase in computation. In practice, however, there are significant algorithmic and performance challenges. In this work, we address these challenges and finally realize the promise of conditional computation, achieving greater than 1000x improvements in model capacity with only minor losses in computational efficiency on modern GPU clusters. We introduce a Sparsely-Gated Mixture-of-Experts layer (MoE), consisting of up to thousands of feed-forward sub-networks. A trainable gating network determines a sparse combination of these experts to use for each example. We apply the MoE to the tasks of language modeling and machine translation, where model capacity is critical for absorbing the vast quantities of knowledge available in the training corpora. We present model architectures in which a MoE with up to 137 billion parameters is applied convolutionally between stacked LSTM layers. On large language modeling and machine translation benchmarks, these models achieve significantly better results than state-of-the-art at lower computational cost."So, here's a big A.I. idea for you: what if we all get one of these sparse Mixture of Experts (MoEs) that's a 100 GB on our SSDs, contains all of the "outrageously large" neural network insights that would otherwise take specialized computers, and is designed to run effectively on a normal GPU or even smaller (e.g. smartphone)?

4 条评论

filterfiber超过 1 年前

评论 #38574779 未加载

dang超过 1 年前

gryfft超过 1 年前

[2017]

评论 #38573015 未加载

评论 #38573502 未加载

legel超过 1 年前

Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts (2017)

4 条评论

Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts (2017)

4 条评论