Mixture of Experts
A model architecture where only a subset of experts activate per token, slashing inference cost while preserving quality.
Text reviewed October 5, 2026
A model architecture where only a subset of experts activate per token, slashing inference cost while preserving quality.
Basic
Instead of running every parameter for every token, MoE routes each token to a small group of specialist sub-networks called experts. DeepSeek V3 uses 671B total parameters but only activates 37B per token. You get frontier quality at a fraction of the compute cost.
Deep
In a dense transformer, every parameter fires for every token. In a Mixture-of-Experts transformer, the feedforward layer is replaced by a bank of experts and a router. For each token the router picks the top-k experts (usually 2 out of 64 or more) and only those experts compute. The trade-offs: load balancing loss to keep experts busy, routing instability during training, and higher total VRAM at serving time because all experts must be resident even if only a few are used per token. See the related terms and live BenchGecko data for current examples.
Expert
MoE replaces the FFN of each transformer block with N expert MLPs plus a gating network g(x) that produces a sparse top-k distribution. The forward pass becomes sum over i in top_k(g(x)) of g(x)_i * E_i(x). Auxiliary losses balance expert load (load_balancing_loss) and prevent gate dropout. Frontier MoE models push expert count to 256+ (DeepSeek V3 uses 256 routed experts + 1 shared expert). Sparsity ratio (active/total params) typically lands between 5% and 15%. Training MoE requires expert parallelism on top of tensor parallelism, and serving benefits from fused kernels like GroupedGEMM. Quality recovery relative to a dense model of the same active size is the empirical win: DeepSeek V3 at 37B active matches Llama 3.1 405B dense on most benchmarks.
Every frontier lab is shipping MoE in 2026.
Depending on why you're here
- ·Like a hospital: you see the specialist for your problem, not every doctor
- ·Models stay smart while becoming cheap to run
- ·MoE models route each token to a subset of experts · output behaves like a normal LLM
- ·All experts must be in VRAM at serve time · requires multi-GPU setups
- ·MoE crashes inference cost · compute-per-dollar dropping ~10× per year
- ·Signals frontier-capability at low price point · disrupts pure API revenue margins
- ·Capex story shifts from more GPUs to smarter architectures
- ·Sparse activation with gated routing, aux loss for load balancing
- ·Scaling law differs from dense: active params matter for quality, total params matter for breadth
- ·See DeepSeek V3 paper (Dec 2024) and Switch Transformer (2021)
Often confused with
Dense = every parameter fires per token. MoE = sparse subset fires per token.
Ensembles combine independent models at inference. MoE is one model with internal routing during forward pass.
MoE is why the AI price floor just dropped by 30×. Any model that isn't MoE by end of 2026 will be priced out of the commodity tier.
Frequently Asked Questions
Read the primary sources
- DeepSeek V3 paperarxiv.org
- Switch Transformer (Google 2021)arxiv.org
- Outrageously Large Neural Networks (original MoE, 2017)arxiv.org