What it is
Most older language models are dense: every part of the network works on every token. A mixture-of-experts model, often shortened to MoE, is built differently. Inside it sit many smaller sub-networks called experts, plus a small router that decides which few experts should handle each token.
Picture a hospital. A dense model is one doctor who treats everything. An MoE model is a hospital with many specialists and a triage nurse at the door. Each patient sees only the specialists they need, so the hospital can employ far more expertise without every patient waiting for every doctor.
Total versus active parameters
This design is why you will see two numbers quoted for the same model.
- Total parameters is the size of the whole network, all experts included.
- Active parameters is how many are actually used for any one token.
A model might have hundreds of billions of total parameters but use only a small fraction at a time. Compute cost and speed during inference track the active number, which is why an MoE model can feel fast and cheap for its size.
The catch
You still have to store the whole model. All experts must sit in memory even though only some run at once, so self-hosting a large MoE model needs serious hardware. Training is also trickier, because the router has to spread work sensibly. If it sends everything to the same few experts, the rest are wasted.
Why it matters to you
If you use hosted AI, MoE is mostly invisible, but it helps explain why newer models can be larger and still affordable. If you run models yourself, particularly an open-weight model, the total-versus-active split decides what hardware you need and what each answer costs.
What it is not
Experts are not neat human subjects like "the law expert" or "the French expert." The router learns its own divisions during training, and they are often hard to describe in plain terms. Also, MoE does not by itself make a model smarter. It makes a given amount of capability cheaper to run.
The window the model can read is separate, set by its context window. Architecture and context length are independent choices.
The takeaway
When you see "X billion parameters, Y billion active," read the active number for speed and cost, and the total number for memory.