Mixture of Experts: how to scale AI models without a proportional rise in costs
Mixture of Experts (MoE) lets AI models grow in capacity without a proportional rise in the cost of every query. How it works, why the largest labs use it, and what it means for performance, infrastructure and cost.

.jpg)
Mixture of Experts (MoE) is one of the architectures that make it possible to build ever larger AI models without a proportional increase in the cost of each query. How does it work, why do the largest AI labs use it, and what does it mean from the perspective of performance, infrastructure and cost?
DeepSeek-V3 has 671 billion parameters. That does not mean, however, that all of them are used when generating each successive token. At any given moment the model activates around 37 billion parameters — less than 6% of its total capacity.
The remaining parameters are still part of the model and remain available, but they do not take direct part in every computation. This is the key difference that makes it possible to substantially increase model capacity without an equally rapid increase in inference cost.
This is not a compromise forced by technological limitations, but a deliberate architectural decision. The mechanism is called Mixture of Experts (MoE), and it is one of the most important ways of decoupling two things that were tightly linked in classic models: the total capacity of the model and the cost of running it.
In practice it answers a very concrete business and engineering question: how do you increase an AI model's capabilities without increasing the cost of every processed query at the same rate?
A hospital as a model of an expert organisation
A good analogy for Mixture of Experts is a large hospital employing many specialists: cardiologists, neurologists, orthopaedists, dermatologists.
A patient is not consulted by the entire medical staff at every visit. Their case is first assessed, and then directed to those specialists whose knowledge is most relevant to the specific problem.
MoE works in a similar way. A router evaluates each fragment of the input — each token — and selects a small group of experts to process it. The remaining experts take no part in that particular computation.
In most modern implementations a token does not go to just one expert. The router may select two, four or more experts at once. This mechanism is known as top-k routing, where k denotes the number of activated experts.
So it is best thought of not as handing the problem to a single specialist, but as a short, automatically assembled consultation with the few most relevant experts.
How Mixture of Experts works under the hood
In the transformer architecture, on which most modern large language models are based, a significant share of the computation takes place in the feed-forward layers.
In models using MoE, a single layer of this type is replaced by a set of parallel sub-networks. Each of them acts as a separate expert.
For each token the process looks, in simplified terms, as follows:
- The router computes a score for how well the token matches each expert.
- The k highest-scoring experts are selected — typically a handful out of a pool of dozens or more.
- Normalised weights are computed for the selected experts, determining their contribution to the result.
- The token is processed by the selected sub-networks, and their outputs are then combined according to the assigned weights.
The order of these operations matters. The occasionally encountered claim that "softmax selects the experts" is a simplification that obscures the essence of the mechanism. In the typical scheme, the top-k experts are selected first, and only afterwards are their scores normalised and used to determine each output's contribution.
From a business and infrastructure point of view, however, the most important thing is the distinction between total parameters and active parameters.
Total parameters describe the model's full capacity. Active parameters indicate what share of that capacity takes part in the computation for a given token.
MoE allows these two figures to be partly decoupled. A model can therefore have hundreds of billions of parameters while using only a small fraction of them during a single inference step.
The naming of the Qwen3-235B-A22B model is a good example. It denotes roughly 235 billion total parameters and 22 billion parameters activated during computation.
There is an important caveat, however: a smaller number of active parameters primarily limits the computational cost of inference, but does not eliminate the cost of storing the entire model.
In a typical server deployment the experts must still be available in memory so that the router can activate them at any moment. For this reason MoE does not remove VRAM requirements. There are techniques for offloading inactive experts to slower RAM or to disk, but this comes at the cost of throughput and response time.
From research concept to production architecture
Mixture of Experts is not a new idea. The concept has been developing for more than three decades, and its current popularity is the result of successive breakthroughs, both algorithmic and infrastructural.
1991 - Adaptive Mixtures of Local Experts
Jacobs, Jordan, Nowlan and Hinton presented the concept of multiple specialised networks combined with a gating mechanism that learned to route particular cases to the appropriate experts.
The scale was incomparable with today's language models, but the basic idea remains very similar: instead of using the whole model for every problem, we dynamically select the most appropriate part of its capacity.
2017 - Outrageously Large Neural Networks
Noam Shazeer and co-authors at Google brought Mixture of Experts into the world of large-scale deep learning.
The paper demonstrated that it was possible to build layers containing a very large number of expert sub-networks and to use sparse routing to activate only a small subset of them. The top-k mechanism introduced there became one of the foundations of later MoE architectures.
2020 - GShard
The next stage was primarily about infrastructure. Designing a huge model is not enough if it cannot be trained efficiently across thousands of accelerators.
GShard showed how to scale MoE models on large TPU infrastructure. It also introduced expert load-balancing mechanisms that limit situations in which the router over-uses a small portion of the available sub-networks.
2021 - Switch Transformer
Switch Transformer simplified routing by limiting the number of activated experts to one per token.
This reduced inter-device communication and simplified computation, while making it possible to experiment with models reaching trillion-parameter scale.
2024 - Mixtral 8x7B
Mixtral was one of the models that significantly widened access to the MoE architecture beyond the largest research labs.
The model uses eight experts, two of which are activated for each token. In total it has around 47 billion parameters, while the active portion of the model is around 13 billion.
It thereby demonstrated the practical benefit of MoE: the ability to achieve quality characteristic of larger models at a significantly lower computational cost per inference.
2024 - DeepSeek-V3
DeepSeek-V3 developed this idea at a far larger scale: 671 billion total parameters with around 37 billion active parameters.
The DeepSeekMoE architecture distinguishes between shared experts, which are activated regardless of routing, and routed experts, responsible for more specialised processing.
DeepSeek also proposed a way of balancing load without the classic auxiliary loss function. From the perspective of MoE's development, this is an important step towards more efficient training of very large sparse models.
It is also worth mentioning Expert Choice Routing, presented by Google in 2022. In this approach the mechanism is reversed: instead of assigning experts to tokens, the experts select the tokens they will process. The goal is primarily better management of individual experts' load.
Which models use Mixture of Experts?
When comparing the largest AI models, it is worth distinguishing data officially published by their creators from estimates prepared by industry analysts.

The status of this data matters. For some closed models the exact architecture is not publicly available, so the figures circulating in the industry should be treated as estimates rather than as parameters confirmed by the vendor.
OpenAI, for example, has not published official confirmation that it uses MoE in its closed models. Information appearing in industry analyses should therefore not be treated on a par with the technical documentation of open models.
Regardless of the details of individual implementations, the trend is clear: as models grow in scale, architectures that allow capacity to increase without a proportional rise in inference cost become increasingly important.
That is precisely the core value of Mixture of Experts.
Advantages and limitations of Mixture of Experts
From the point of view of organisations developing or deploying AI solutions, MoE offers several significant benefits.
Greater capacity without a proportional rise in compute cost
A model can have far more parameters than it uses during a single inference step. This makes it possible to increase its capacity without having to perform the full set of computations for every token.
More efficient inference
Activating only a selected part of the network makes it possible to obtain the properties of a larger model with fewer operations performed while generating a response.
Automatic expert specialisation
Individual sub-networks can specialise during training in particular types of data or problems — for example code, mathematics or specific languages. This division does not have to be defined manually by the model's designers.
Better economics of training at scale
Research on Switch Transformer and subsequent architectures has shown that sparse MoE can significantly improve the training efficiency of very large models within a given compute budget.
This architecture also comes, however, with additional costs and complications.
High memory requirements
The number of active parameters may be relatively small, but the full set of experts must still be stored and available to the system. From an infrastructure perspective this means memory demand can remain very high.
Greater infrastructure complexity
Experts may be distributed across multiple accelerators or nodes. This requires managing communication and synchronisation within so-called expert parallelism, which complicates serving compared with a classic dense model.
The load-balancing problem
The router may start to favour a small group of experts. This leads both to uneven use of infrastructure and to a situation in which some experts are trained far less often than others.
For this reason MoE architectures employ additional mechanisms to ensure a more even distribution of traffic.
Greater fine-tuning complexity
Fine-tuning MoE models has historically involved problems with routing and balance between experts. The development of techniques such as MoE-LoRA has significantly reduced this problem, but it remains an additional element that must be taken into account when designing the fine-tuning process.
Training stability
In the first large MoE implementations, training stability was one of the major challenges. Subsequent generations of models introduced mechanisms that significantly improved the situation, so today this is more an engineering problem requiring appropriate design than a fundamental limitation of the architecture itself.
What does Mixture of Experts mean in practice?
Mixture of Experts is no longer merely an interesting research solution. It is one of the AI industry's key answers to the problem of the rising cost of scaling models.
For anyone evaluating a model, the most important concept should be the distinction between its total capacity and the portion activated during inference.
A designation such as 235B-A22B therefore conveys more information than a parameter count alone. It says not only how large the model is, but also what share of its resources takes part in the computation for a single token.
This has direct implications for deployment economics. The total parameter count alone no longer allows reliable conclusions about the cost of using a model. You also have to take into account the architecture, the number of active parameters, memory requirements and how experts are distributed across devices.
For teams choosing models for production use, this means that comparing models purely on parameter count is becoming less and less useful.
Far more important are the questions:
- what share of the model is activated during inference,
- what the actual cost of processing tokens is,
- what memory resources the deployment requires,
- what throughput the solution offers,
- and what quality the model delivers at a given infrastructure cost.
This is exactly where Mixture of Experts shows its greatest value.
It allows model capacity to grow without performing every possible computation each time. Instead of scaling cost together with the parameter count, the system dynamically selects those parts of the model that are most needed to solve a specific problem.
It is largely thanks to this approach that successive generations of models can grow far faster than the cost of a single query.
Sources
- Shazeer et al., Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer (2017)
- Lepikhin et al., GShard: Scaling Giant Models with Conditional Computation and Automatic Sharding (2020)
- Fedus, Zoph, Shazeer, Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity (2021)
- Jiang et al., Mixtral of Experts (2024)
- DeepSeek-AI, DeepSeek-V3 Technical Report (2024)
- Jacobs, Jordan, Nowlan, Hinton, Adaptive Mixtures of Local Experts (1991)
- Zhou et al., Mixture-of-Experts with Expert Choice Routing (2022)
