Mixture of experts
Transcript: cleaned auto-captions with timestamps
Lecture 3 walked the space of dense Transformer variants and concluded that the architecture has largely converged. Lecture 4 is the exception clause: the one architectural change since 2023 that actually moved the frontier, and the one every serious open model made in 2024–2025. Hashimoto opens by saying this used to be a bonus lecture and is now a core one — which is itself the argument. If you want the best model per FLOP you can afford, you build a sparse one, and this lecture is where the course tells you what that costs in optimization pain and systems complexity before the systems arc (lectures 5–8) makes you pay it.
Outline, with timestamps
- 00:05 — Why this became a required lecture: GPT-4's rumored MoE shape, Grok, DeepSeek, Llama 4.
- 01:43 — What an MoE actually is: split the FFN into N copies, add a router, activate k. The name is misleading.
- 03:52 — The evidence: FLOP-matched loss keeps dropping as experts are added (Switch Transformer, OLMoE, DeepSeek-V2).
- 08:12 — Expert parallelism: experts are a natural sharding axis, and multi-node is where MoEs win biggest.
- 12:33 — Why MoEs stayed niche: infrastructure complexity and a non-differentiable, sometimes unstable objective.
- 15:51 — Three routing families: token choice, expert choice, global assignment — and why one won.
- 21:14 — Top-k routing in detail: the router equation, where the softmax goes, why k = 2 became canonical.
- 22:22 — The uncomfortable baseline: hash routing works, and RL / linear-assignment routing mostly does not pay.
- 31:52 — Expert geometry: fine-grained experts, shared experts, and the configuration table across a dozen models.
- 41:47 — Training a router you cannot differentiate: RL, stochastic perturbation, or heuristic balancing losses.
- 48:22 — The load-balancing loss everyone uses, and DeepSeek-V3's bias-only replacement for it.
- 59:54 — Systems: all-to-all dispatch and combine, block-sparse kernels, token dropping, batch-level nondeterminism.
- 64:48 — Stability and fine-tuning: fp32 routers, router z-loss, sparse-model overfitting.
- 68:43 — Upcycling: initialize an MoE from a trained dense model (MiniCPM, Qwen1.5-MoE).
- 70:20 — Worked example: DeepSeekMoE v1 → V2 → V3, plus the non-MoE parts (MLA, MTP).
The trade: parameters are cheap, FLOPs are not 03:19
Strip the marketing off and an MoE layer is one substitution. Wherever a dense Transformer block has a single feed-forward network, you put N feed-forward networks and a router that picks k of them for each token's hidden state. Everything else — attention, norms, residual stream, embeddings — is untouched. The name invites you to imagine a coding expert and a French expert; that is not what you get, and Hashimoto spends his first two minutes killing that intuition:
"I think you hear mixture of experts and you think, oh, there must be experts specialized for different domains… it is very far from that mental model."— Tatsunori Hashimoto, 01:43
The arithmetic is what matters. If each expert is the size of the dense FFN you replaced and you activate k = 1, the forward pass does exactly the same matrix multiplies as the dense model — same FLOPs — while the model holds N times the FFN parameters. Parameters are memory and memory is comparatively cheap; FLOPs are the training budget. So the question is entirely empirical: does the extra memorization capacity convert into loss?
It does, repeatedly. The Switch Transformer line of work shows FLOP-matched training loss falling monotonically as you go from 1 to 128 to 256 experts, and the 128-expert Switch-Base reaching a target perplexity roughly 7× faster in steps than its dense twin. That is a 2021–2022 result, so the fair objection is that it might not survive modern recipes; OLMoE re-ran the comparison in 2024 with careful controls and got the same picture 04:56. The vendor version of this plot is DeepSeek-V2's MMLU-versus-activated-parameters chart, which Hashimoto flags for what it is — a little sleight of hand, since the x-axis silently drops every parameter you still had to store and shard 07:07. It is still the right axis if inference FLOPs are your binding constraint, which for a deployed model they usually are.
The second reason people build MoEs is parallelism, not loss. Experts are a natural sharding boundary — put expert i on device i — so expert parallelism becomes an axis you compose with data and tensor parallelism, and the advantage is largest exactly where you were going to shard anyway 08:12. Below multi-node scale a dense model is simply less trouble, which is much of why this was never the default thing taught in an NLP class.
Routing: three families, and the boring one won 15:51
Think of a routing scheme as a scoring matrix with tokens on one axis and experts on the other. You can take the top k down each token's column (token choice), the top k across each expert's row (expert choice), or solve a global assignment problem that balances both (optimization-based routing). Expert choice has an attractive property for free — every expert gets exactly the same number of tokens, so device utilization is balanced by construction. Global assignment is the elegant answer. And essentially every shipped model does token choice top-k, because it is the one that scores tokens by "which expert will process me well" rather than by scheduling convenience, and because OLMoE's ablations show token choice descending noticeably faster in validation loss 16:57.
The router itself is smaller than people expect. Each expert i owns a learned vector e_i; for a hidden state u you compute the inner products u · e_i, normalize them, keep the k largest, and use those normalized scores as gate weights on a weighted sum of the k expert outputs, which is then added back to the residual stream. That is the whole thing — structurally an attention score against a small set of learned keys. The only variation across models is where the normalizer sits: DeepSeek v1–V2, Grok and Qwen softmax the affinities before the top-k, while Mixtral, DBRX and DeepSeek-V3 take the top-k first and normalize after 24:03. Hashimoto's read is that this is close to aesthetic: nothing requires the gates to sum to one, because the next layer norm can rescale whatever you hand it 26:22. The softmax here is a normalize-to-one operation, not a soft argmax.
Two questions from the room are worth carrying away. Why keep the top-k at all instead of gating all N experts softly? Because the sparsity is the product — dense gating pays the FLOPs of all N experts at training time and the premise evaporates 27:30. And why is the router a single linear map? Partly because router FLOPs come out of your budget, but mainly because there is so little signal to learn a better router with: your only information about expert quality comes from the k experts you actually ran, so a richer parameterization has nothing extra to fit 30:48.
Which brings up the result that should make you suspicious of the whole enterprise: hash routing works. Replace the learned router with a fixed hash of the token and you still get most of the gains over dense 22:22. Hashimoto's explanation is that hashing is deterministic, so an expert still sees a consistent slice of the input distribution and can specialize on it — non-semantically, and with Zipfian token frequencies possibly semantically after all, since a single expert may end up owning "the" 56:37. The prediction he offers, untested, is that input-independent random routing would be terrible. The gap between hashing and learned routing is roughly a measurement of how much your router is actually earning.
The more principled alternatives have been tried and are not used. Reinforcement learning on the routing policy is the textbook-correct answer to a discrete decision and appears in the earliest conditional-computation work; Clark et al.'s scaling study includes an RL-R baseline that loses to their Sinkhorn-based S-BASE assignment method, and gradient variance plus complexity have kept it out of production 22:56. Linear-assignment and optimal-transport routers are elegant and cost more than they return 23:29.
Expert geometry: fine-grained and shared 31:52
Given that more experts help, the obvious follow-up is: do they have to be full-size? DeepSeekMoE's answer is no — shrink each expert's hidden dimension by a factor of r and you can afford r times as many of them at the same parameter count, then raise k proportionally so activated FLOPs are unchanged. If the dense rule of thumb is d_ff ≈ 4·d_model (or ≈ 2.6·d_model for a gated MLP), a "1/4 fine-grained" expert simply uses a quarter of that projection width. This is close to free combinatorics: with 64 quarter-size experts and 6 active you get vastly more routing configurations than 16 full-size experts with 2 active, at the same cost 32:57. Both DeepSeek's own ablations and OLMoE's independent replication (8 → 32 → 64 fine-grained experts) show clean monotone gains.
The companion idea is the shared expert: one or two MLPs that every token goes through unconditionally, on the theory that some processing is common to all tokens and it is wasteful to make the router rediscover that, and to replicate it across every routed expert 34:02. Here the evidence genuinely splits. DeepSeek reports a solid boost; OLMoE's controlled comparison finds essentially nothing and ships without shared experts 36:19. This is exactly the kind of divergence worth holding loosely — fine-grained experts are a no-brainer, shared experts are a coin whose bias nobody has pinned down. Note the field's drift: many shared experts (Qwen 1.5's four) gave way to one or zero.
| Model | Routed | Active | Shared | Fine-grained ratio |
|---|---|---|---|---|
| GShard | 2048 | 2 | 0 | — |
| Switch Transformer | 64 | 1 | 0 | — |
| ST-MoE | 64 | 2 | 0 | — |
| Mixtral | 8 | 2 | 0 | — |
| DBRX | 16 | 4 | 0 | — |
| Grok | 8 | 2 | 0 | — |
| DeepSeekMoE (v1) | 64 | 6 | 2 | 1/4 |
| Qwen 1.5-MoE | 60 | 4 | 4 | 1/8 |
| DeepSeek-V3 | 256 | 8 | 1 | 1/14 |
| OLMoE | 64 | 8 | 0 | 1/8 |
| MiniMax | 32 | 2 | 0 | ~1/4 |
| Llama 4 Maverick | 128 | 1 | 1 | 1/2 |
From the lecture slides 37:23. Hashimoto flags the ratio column as back-derived from config files rather than reported, so treat it as approximate.
Read the table top to bottom and the trajectory is legible: an early Google phase of very many full-size experts, a 2023–2024 Western phase of 8–16 full-size experts with k = 2, and then everyone converging on the DeepSeekMoE shape — many small experts, more of them active, zero or one shared. On k itself, the original argument for k ≥ 2 was exploration: with one active expert you only ever exploit, whereas two lets the gradient compare them 19:37. Once experts are quarter-size, raising k is FLOP-neutral, which is why fine-grained models happily run k = 6 or 8.
Training a router you cannot differentiate 42:19
This is the actual content of the lecture. Sparsity is mandatory at training time, top-k is a hard discrete selection, and hard selections have no gradient. Three families of fixes exist, and the field picked the least respectable one.
RL is the principled option and is not used, for the reasons above. Stochastic perturbation is the middle option: Shazeer et al.'s noisy top-k gating adds learned-scale Gaussian noise to the router logits before selecting, which turns routing into an epsilon-greedy bandit — you occasionally pull an arm you would not have, which both explores and produces experts that are less brittle because they see tokens they did not expect 44:30. The Switch Transformer's multiplicative "jitter" is the same idea in a cheaper form. Both were largely abandoned — Zoph et al. removed jitter in ST-MoE — because they cost specialization and the loss-based approach worked better 46:11.
What everyone actually does is heuristic load balancing. Left alone, top-k routing has a brutal failure mode: early in training one expert is marginally better, the router sends it more tokens, it improves faster, and you converge to a local minimum where one or two experts do everything and the rest are dead weight you are still paying memory for 47:50. OLMoE's ablation makes this vivid — remove the balancing loss and two of eight experts absorb roughly half the tokens while the rest go dark 58:48. The auxiliary loss from the Switch Transformer, which almost every MoE since has copied, is:
L_aux = α · N · Σ_i f_i · p_i
where f_i is the fraction of tokens actually dispatched to expert i over the balancing unit, and p_i is the fraction of router probability mass intended for expert i. The dot product is minimized when both vectors are uniform. The instructive part is the derivative with respect to p_i, which is proportional to f_i: an expert that received more tokens gets pushed down harder, so the loss self-corrects in proportion to how imbalanced things already are 50:02. Note that f_i is a hard count and therefore has no useful gradient itself — the learning signal flows entirely through p_i, which is why the loss is written as a product of the two rather than as a divergence between them.
The choice of balancing unit is where variants live. Per-expert-per-batch is the default. DeepSeek adds a per-device version of the identical formula, summing token counts over device groups rather than individual experts, so the objective directly optimizes GPU utilization 51:07. DeepSeek-V3 then does something genuinely new — the first MoE training idea in the lecture that did not come out of Google. It deletes the per-expert auxiliary loss entirely and replaces it with a per-expert bias b_i added to the affinity score for selection purposes only, never propagated into the gate weight. b_i is updated by online gradient-free rules: after each batch, add γ to underloaded experts and subtract γ from overloaded ones 52:12. Because it never touches the gate, it introduces no interference gradient into the language-modeling objective — that is the actual selling point, and DeepSeek calls it "auxiliary-loss-free load balancing." Hashimoto's footnote is the memorable part:
"So it's not fully auxiliary-loss-free as they'd like you to believe."— Tatsunori Hashimoto, 54:26
They kept a complementary sequence-wise auxiliary loss, and the reason is worth understanding rather than reading as hypocrisy. Batch-level balance is fine during training when you control the batch. At inference you do not: one user can send a wildly out-of-distribution sequence that overwhelms a handful of experts. Sequence-level balancing is insurance against a serving-time pathology, not a training-time one 75:46.
Where the systems bill comes due 59:54
The runtime shape of an expert-parallel MoE layer is: router → all-to-all dispatch shipping each token's hidden state to the devices holding its chosen experts → local FFN compute → all-to-all combine bringing results home to be summed into the residual stream 60:27. Two collectives per MoE layer per forward pass, plus their mirrors in the backward pass — affordable only if the expert compute is big enough to amortize them, which is the real reason fine-grained experts cannot be pushed arbitrarily far.
When several experts share a device, the per-expert matmuls are small and ragged, and naive implementations either pad to a fixed capacity (wasting FLOPs) or drop the overflow (losing tokens). MegaBlocks reformulates the layer as one block-sparse matrix multiply with custom kernels and avoids both 61:31. It sits under a lot of open MoE training.
The token-dropping story has a delightful consequence. Systems that do cap per-expert load use a capacity factor; when a batch overloads an expert, the excess tokens are simply skipped — the MLP contributes nothing and the residual connection passes the hidden state through unchanged. Because capacity is enforced per batch, whether your token gets dropped depends on who else is in your batch. This is a real source of nondeterminism at temperature zero, and was one of the popular explanations for GPT-4's early nondeterministic behavior. Hashimoto is careful not to assert it was the cause, only that the mechanism exists and that cross-batch coupling at inference is something engineers rarely think about 63:43.
Stability, fine-tuning, and upcycling 64:48
Lecture 3's stability rule was "fear the softmax," and the router is a softmax over a handful of logits inside an otherwise bf16 model. ST-MoE's two prescriptions are cheap and now standard: compute the router in float32 regardless of ambient precision, and add a router z-loss — the squared log-sum-exp of the router logits — to hold the normalizer near one 65:57. Ablate the z-loss and the validation curve grows large spikes; training recovers each time but ends measurably worse 66:30. It is the same z-loss that later became routine on the main LM head.
Fine-tuning is the other soft spot: sparse models carry far more parameters than their FLOPs suggest, so on small supervised sets they overfit hard, with a train/val gap the dense baseline does not show 67:05. ST-MoE proposed interleaving dense and MoE layers and fine-tuning only the dense MLPs; Hashimoto's read is that it never caught on. DeepSeek's unglamorous fix did: use enough data — 1.4M SFT examples 68:10.
Upcycling is the most practical trick here if you already own a dense model: copy its MLP N times, perturb the copies, initialize a fresh router, keep training 68:43. MiniCPM demonstrated it with roughly 520B additional tokens; Qwen1.5-MoE upcycled a 1.8B dense model into a 2.7B-active MoE that reached parity with their 7B dense line 69:48.
DeepSeek v1 → V2 → V3, as the worked example 70:20
The closing walkthrough exists to make a point about how little architecture actually changed across a 40× scale-up. DeepSeekMoE (v1), 16B total / 2.8B active: 2 shared plus 64 quarter-size fine-grained experts, standard softmax-then-top-k routing, standard auxiliary balancing at expert and device level. DeepSeek-V2, 236B / 21B active: literally the same MoE diagram with different counts, plus two systems-driven additions — top-M device routing, which first restricts a token to its M best devices and only then takes top-k within them (a direct cap on how fragmented the all-to-all can get), and a communication balancing loss that balances the outbound leg, not just the inbound one 72:33. DeepSeek-V3, 671B / 37B active: 1 shared plus 256 fine-grained experts with 8 active, sigmoid affinities with normalization moved after the top-k, top-M device routing retained, communication loss dropped, and the aux-loss-free bias scheme plus sequence-wise loss described above.
"The MoE architecture itself doesn't change. That's stayed the same since DeepSeekMoE. Like, if it works, don't change it."— Tatsunori Hashimoto, 74:42
The last ten minutes cover the two non-MoE pieces of V3, because at that point they are all that remains unexplained. MLA (multi-head latent attention) is a KV-cache compression that takes a different route than GQA/MQA from lecture 3: instead of reducing the number of KV heads, project the hidden state down to a small latent c and cache that, up-projecting to K and V on use. The apparent extra matmul is free because W_UK can be folded into the query projection by associativity — you never materialize the up-projection separately 78:31. The catch is RoPE: the rotation matrices sit between the query projection and the up-projection, and matrix products do not commute, so the folding trick breaks. DeepSeek's answer is to keep a few uncompressed key dimensions that carry the rotation 79:35. MTP (multi-token prediction) adds a lightweight one-layer transformer head on top of the hidden state that predicts a second future token, giving denser training signal (and a natural speculative-decoding draft head, in the EAGLE tradition). Hashimoto notes with some disappointment that despite a diagram implying depth D, V3 only ever predicts one token ahead 80:08.
What you build with this
No assignment starts here. You are still inside Assignment 1: Basics (handout, leaderboard), which builds a dense Transformer end to end — BPE tokenizer, RMSNorm, SwiGLU FFN, RoPE, multi-head attention, AdamW with cosine schedule and gradient clipping, and a training loop tuned against a fixed compute budget. Nothing in this lecture is required for it. If you want to use it, the leaderboard's fixed-FLOP framing is precisely the setting where a sparse FFN should win, but you would be hand-rolling the routing and the balancing loss with no kernel support, and single-GPU MoE gives up the parallelism half of the payoff — a research detour, not a shortcut.
Where the lecture pays off is Assignment 2: Systems (handout), opened by L05 GPUs and continued through kernels and the two parallelism lectures. Every cost this lecture waves at — the all-to-all dispatch and combine, why a ragged per-expert matmul needs block-sparse kernels, why capacity factors and token dropping exist, why balancing across devices is a term in the loss — becomes concrete once you have benchmarked collectives and written a kernel yourself. Read this lecture as the motivating workload for that arc, and revisit its systems section after L07.
Supporting materials, verified
- A Review of Sparse Expert Models in Deep Learning — Fedus, Dean, Zoph (2022) · The survey most of the lecture's figures come from; the routing taxonomy (token choice / expert choice / global assignment) is straight out of it.
- Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity — Fedus, Zoph, Shazeer (2021) · Source of the k = 1 router, the f·p load-balancing loss everyone still uses, and the FLOP-matched scaling plots.
- Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer — Shazeer et al. (2017) · The noisy top-k gate with a learned noise scale, i.e. the stochastic-perturbation branch of the design space.
- GShard: Scaling Giant Models with Conditional Computation and Automatic Sharding — Lepikhin et al. (2020) · The k = 2 / 2048-expert baseline configuration in the comparison table, and the origin of expert-parallel sharding annotations.
- ST-MoE: Designing Stable and Transferable Sparse Expert Models — Zoph et al. (2022) · The stability lecture-within-the-lecture: fp32 routers, the router z-loss, the removal of jitter, and the sparse fine-tuning overfitting results.
- DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language Models — Dai et al. (2024) · Introduces fine-grained expert segmentation plus shared expert isolation, with the ablations Hashimoto walks through; the v1 in the closing v1→V3 narrative.
- DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model — DeepSeek-AI (2024) · Top-M device-limited routing, the communication balancing loss, and the original MLA formulation.
- DeepSeek-V3 Technical Report — DeepSeek-AI (2024) · The lecture's closing worked example: sigmoid gating, 256 routed + 1 shared experts, bias-based balancing, sequence-wise auxiliary loss, MLA, MTP.
- Auxiliary-Loss-Free Load Balancing Strategy for Mixture-of-Experts — Wang et al. (2024) · The standalone paper behind V3's per-expert bias update rule, with the ablations justifying it over an auxiliary loss.
- OLMoE: Open Mixture-of-Experts Language Models — Muennighoff et al. (2024) · The ablation suite the lecture leans on hardest: token vs expert choice, fine-grained expert counts, shared experts (no gain), and the dead-expert plot when balancing is removed.
- Unified Scaling Laws for Routed Language Models — Clark et al. (2022) · The head-to-head of RL-R, HASH layers and Sinkhorn-BASE routing at scale; the source of "RL is principled but doesn't win."
- MegaBlocks: Efficient Sparse Training with Mixture-of-Experts — Gale, Narayanan, Young, Zaharia (2022) · Block-sparse MoE kernels that eliminate both token dropping and padding waste; the library under many open MoE training stacks.
- Mixtral of Experts — Jiang et al. (2024) · The 8-experts / k = 2 Western open model in the table, and an example of normalizing after the top-k.
- DeepSpeed-MoE: Advancing Mixture-of-Experts Inference and Training to Power Next-Generation AI Scale — Rajbhandari et al. (2022) · Where the shared-expert idea originates, per the slides, before DeepSeek and Qwen adopted it.
- Qwen1.5-MoE-A2.7B: Matching 7B Model Performance with 1/3 Activated Parameters — Qwen Team (2024) · The upcycling success story from a 1.8B dense checkpoint; 60 routed + 4 shared experts.
- MiniCPM: Unveiling the Potential of Small Language Models with Scalable Training Strategies — Hu et al. (2024) · The other upcycling example, ~520B additional tokens on top of a dense base.
- Hash Layers For Large Sparse Models — Roller et al. (2021) · field map extra — the fixed-hash router that Hashimoto cites as "you don't even need a smart router," worth reading for how strong the baseline really is.
- BASE Layers: Simplifying Training of Large, Sparse Models — Lewis et al. (2021) · field map extra — the linear-assignment routing the lecture mentions in passing, and the ancestor of Clark et al.'s S-BASE.
- Estimating or Propagating Gradients Through Stochastic Neurons for Conditional Computation — Bengio, Léonard, Courville (2013) · field map extra — the original framing of the non-differentiable-gate problem this lecture spends half its time on.
- JetMoE: Reaching Llama2 Performance with 0.1M Dollars — Shen et al. (2024) · field map extra — one of the rare models that puts MoE on the attention heads as well as the MLPs, the "less common" branch on the slides.
Exercises
- Router collapse in miniature code — Reproduce the dead-expert failure on a toy problem, then fix it. Steps: (1) build a small MoE layer in PyTorch — 8 experts, a linear router, softmax-then-top-2 gating, weighted sum into a residual; (2) train it on any tiny sequence task (character-level LM on a few MB is plenty) with no balancing loss, logging per-expert token counts each step; (3) plot expert load over training and confirm a couple of experts absorb most of the mass; (4) add the Switch auxiliary loss α·N·Σ f_i·p_i with α ≈ 0.01 and re-run; (5) plot load and validation loss for both runs. A good answer shows the collapse curve, shows balance restored, and states whether validation loss improved and by how much — the honest result is sometimes that the loss barely moves at this scale while utilization changes dramatically, which is itself the lesson about what balancing is for.
- Fine-grained expert accounting — Take a dense block with d_model = 1024 and a gated MLP at d_ff ≈ 2.6·d_model. Work out, on paper: (a) the FFN parameter count; (b) the config for 1/4-size fine-grained experts with 64 routed experts and k = 6, giving total and activated parameters; (c) the same for 1/8-size with 128 experts, keeping activated FLOPs matched. Then answer the question the lecture leaves open — what actually stops you at 1/32 or 1/64? A good answer names all three limits: matmuls become too small to saturate the GPU, all-to-all message count grows with k, and router logit count grows with N while the gradient signal per expert shrinks.
- Read the aux-loss-free claim adversarially — Read section 2.1.2 of the DeepSeek-V3 report alongside Wang et al. (2024). Write half a page answering: what exactly does the bias b_i touch and what does it deliberately not touch, and why does that distinction matter for the LM gradient? Then explain why a sequence-wise auxiliary loss was still needed, and why that is a serving concern rather than a training one. A good answer states that b_i affects selection but never the gate weight, so no balancing gradient enters the main objective, and connects the sequence-wise term to inference-time out-of-distribution inputs you cannot batch-average away.
- Upcycle a small dense model code — Take any small pretrained dense LM you can fit. Steps: (1) replicate each MLP block 4× and add small Gaussian perturbations to break symmetry; (2) attach a freshly initialized linear router with top-2 gating and the Switch balancing loss; (3) continue training on a modest token budget, tracking loss against a dense continued-training control on the same budget; (4) log expert load and report whether the router differentiated the copies at all. A good answer reports the crossover point (or its absence) and inspects whether experts specialize by token identity, frequency, or not at all — connecting back to why hash routing is such a strong baseline.