Skip to content
LiveNext 11:35:27

Definition

What is Mixture of experts?

Mixture of experts is a mixture of experts is a model design that splits a neural network into many specialist sub-networks and activates only a few of them for each token, cutting the compute needed per answer.

What it is

Most older language models are dense: every part of the network works on every token. A mixture-of-experts model, often shortened to MoE, is built differently. Inside it sit many smaller sub-networks called experts, plus a small router that decides which few experts should handle each token.

Picture a hospital. A dense model is one doctor who treats everything. An MoE model is a hospital with many specialists and a triage nurse at the door. Each patient sees only the specialists they need, so the hospital can employ far more expertise without every patient waiting for every doctor.

Total versus active parameters

This design is why you will see two numbers quoted for the same model.

  • Total parameters is the size of the whole network, all experts included.
  • Active parameters is how many are actually used for any one token.

A model might have hundreds of billions of total parameters but use only a small fraction at a time. Compute cost and speed during inference track the active number, which is why an MoE model can feel fast and cheap for its size.

The catch

You still have to store the whole model. All experts must sit in memory even though only some run at once, so self-hosting a large MoE model needs serious hardware. Training is also trickier, because the router has to spread work sensibly. If it sends everything to the same few experts, the rest are wasted.

Why it matters to you

If you use hosted AI, MoE is mostly invisible, but it helps explain why newer models can be larger and still affordable. If you run models yourself, particularly an open-weight model, the total-versus-active split decides what hardware you need and what each answer costs.

What it is not

Experts are not neat human subjects like "the law expert" or "the French expert." The router learns its own divisions during training, and they are often hard to describe in plain terms. Also, MoE does not by itself make a model smarter. It makes a given amount of capability cheaper to run.

The window the model can read is separate, set by its context window. Architecture and context length are independent choices.

The takeaway

When you see "X billion parameters, Y billion active," read the active number for speed and cost, and the total number for memory.

Questions people ask

What does active parameters mean?

It is the number of a model's parameters actually used to process each token. In a mixture of experts it is much smaller than the total.

Are mixture-of-experts models cheaper to run?

Per token, usually yes, because only some experts work. But all experts must still be loaded in memory, so hosting needs a lot of it.

Mixture of experts in the news