MoE Dynamic Routing & Expert Parallelism Simulator
Explore how distributed Mixture-of-Experts clusters route tokens dynamically across 16 to 64 fine-grained experts hosted on 8 GPU ranks. Observe Top-2 gating softmax, simulate buffer capacity factor limits ($C$), compare traditional auxiliary loss with DeepSeek-V3 auxiliary-loss-free dynamic bias, and inject live chaos to witness distributed All-to-All straggler barrier bottlenecks in real time.
Mixtral (16E) vs DeepSeek-V3 fine-grained clustering (64E).
Capacity_E = ⌈(T · K / E) · C⌉. Below 1.25x drops tokens during hotspots.
Mathematical regularization technique to avoid single-expert bottleneck.
L_balance = α · K · ∑ f_i · P_i (penalizes gate collapse)
4.00 bits (max: uniform 100%)
14 / 64 slots dropped
All-to-All synchronization wait time (max_i τ_i)
50•Drops: 14Distributed Expert Parallelism Topology (8 GPU Ranks)
All-to-All collective dispatch routes tokens across 8 physical GPU devices. Click any expert or GPU node to inject chaos.
Microbatch Token Dispatch Matrix (Top-2 Gating Softmax)
Per-token routing decisions, normalized gating weights (g_1 + g_2 = 1.0), and expert buffer drop statuses.
| # | Token Vector & Text | Category | Top-1 Expert & Weight (g₁) | Top-2 Expert & Weight (g₂) | Decision Entropy | Dispatch Status |
|---|---|---|---|---|---|---|
| 1 | class TransformerBlock | code | Expert-02(G1)50.8% | Expert-01(G0)49.2% | 95% | ROUTED (2/2) |
| 2 | curriculum distillation | natural | Expert-13(G6)53.4% | Expert-12(G6)46.6% | 95% | ROUTED (2/2) |
| 3 | cudaMemcpyAsync | system | Expert-09(G4)50.8% | Expert-10(G5)49.2% | 94% | ROUTED (2/2) |
| 4 | self-reflective critique | natural | Expert-13(G6)53.3% | Expert-12(G6)46.7% | 95% | ROUTED (2/2) |
| 5 | tensor.contiguous() | code | Expert-02(G1)50.8% | Expert-01(G0)49.2% | 96% | ROUTED (2/2) |
| 6 | in-context few-shot | natural | Expert-13(G6)52.4% | Expert-14(G7)47.6% | 95% | ROUTED (2/2) |
| 7 | ::std::sync::Arc | punctuation | Expert-00(G0)50.2% | Expert-15(G7)49.8% | 98% | ROUTED (2/2) |
| 8 | in-context few-shot | natural | Expert-13(G6)53.0% | Expert-14(G7)47.0% | 95% | ROUTED (2/2) |
| 9 | [0, 1, 2, ...tail] | punctuation | Expert-15(G7)50.7% | Expert-00(G0)49.3% | 98% | ROUTED (2/2) |
| 10 | epoll_wait(ring_fd) | system | Expert-10(G5)51.2% | Expert-09(G4)48.8% | 94% | ROUTED (2/2) |
| 11 | ∇_θ L_total | math | Expert-05(G2)51.2% | Expert-06(G3)48.8% | 94% | ROUTED (2/2) |
| 12 | posix_memalign | system | Expert-10(G5)50.7% | Expert-09(G4)49.3% | 95% | ROUTED (2/2) |
| 13 | optimizer.step() | code | Expert-02(G1)50.2% | Expert-01(G0)49.8% | 96% | ROUTED (2/2) |
| 14 | SRAM_TILE_LOAD_256K | system | Expert-09(G4)50.7% | Expert-10(G5)49.3% | 94% | ROUTED (2/2) |
| 15 | class TransformerBlock | code | Expert-02(G1)50.5% | Expert-01(G0)49.5% | 95% | ROUTED (2/2) |
| 16 | -> Result<T, E> | punctuation | Expert-15(G7)51.4% | Expert-00(G0)48.6% | 98% | ROUTED (2/2) |
| 17 | ∫_{-∞}^∞ e^{-x²} dx | math | Expert-05(G2)50.5% | Expert-06(G3)49.5% | 94% | ROUTED (2/2) |
| 18 | in-context few-shot | natural | Expert-13(G6)53.2% | Expert-14(G7)46.8% | 95% | ROUTED (2/2) |
| 19 | ∫_{-∞}^∞ e^{-x²} dx | math | Expert-05(G2)50.0% | Expert-06(G3)50.0% | 93% | ROUTED (2/2) |
| 20 | optimizer.step() | code | Expert-02(G1)50.3% | Expert-01(G0)49.7% | 95% | ROUTED (2/2) |
| 21 | posix_memalign | system | Expert-10(G5)50.8% | Expert-09(G4)49.2% | 94% | ROUTED (2/2) |
| 22 | eigen(W_gate) | math | Expert-05(G2)50.7% | Expert-06(G3)49.3% | 94% | ROUTED (2/2) |
| 23 | curriculum distillation | natural | Expert-13(G6)52.9%[DROP] | Expert-14(G7)47.1% | 95% | CAPACITY EXCEEDED |
| 24 | self-reflective critique | natural | Expert-13(G6)53.4%[DROP] | Expert-14(G7)46.6% | 95% | CAPACITY EXCEEDED |
| 25 | multi-hop synthesis | natural | Expert-13(G6)53.4%[DROP] | Expert-12(G6)46.6% | 95% | CAPACITY EXCEEDED |
| 26 | eigen(W_gate) | math | Expert-05(G2)50.4% | Expert-06(G3)49.6% | 94% | ROUTED (2/2) |
| 27 | async function eval | code | Expert-01(G0)51.1%[DROP] | Expert-02(G1)48.9%[DROP] | 95% | CAPACITY EXCEEDED |
| 28 | tensor.contiguous() | code | Expert-02(G1)51.1%[DROP] | Expert-01(G0)48.9%[DROP] | 96% | CAPACITY EXCEEDED |
| 29 | ∫_{-∞}^∞ e^{-x²} dx | math | Expert-05(G2)50.5%[DROP] | Expert-06(G3)49.5%[DROP] | 94% | CAPACITY EXCEEDED |
| 30 | self-reflective critique | natural | Expert-13(G6)53.3%[DROP] | Expert-12(G6)46.7% | 95% | CAPACITY EXCEEDED |
| 31 | cudaMemcpyAsync | system | Expert-09(G4)50.6%[DROP] | Expert-10(G5)49.4%[DROP] | 94% | CAPACITY EXCEEDED |
| 32 | All-to-All Dispatch | system | Expert-09(G4)50.7%[DROP] | Expert-10(G5)49.3%[DROP] | 94% | CAPACITY EXCEEDED |
MoE Router & Distributed Collective Event Stream
Auxiliary Loss vs DeepSeek-V3 Dynamic Bias
Switch between Aux Loss (L_balance = α · K · ∑ f_i · P_i) and DeepSeek Bias (b_i ← b_i + γ(1/E - f_i)). Observe how DeepSeek-V3 autonomously steers traffic away from congested experts without injecting surrogate gradient penalties into the core representation loss.
Capacity Factor (C) & Buffer Waste
Modulate the Capacity Factor slider from 1.0× to 2.0×. At C=1.0× (zero headroom), even slight Poisson variance causes token drops. At C=2.0×, drops fall to 0%, but idle padding slots consume up to 50% of allocated GPU tensor buffers.
Celebrity Hotspot Skew & Straggler Barrier
Click Celebrity Hotspot (+15.0). Watch tokens collapse into hot experts, causing entropy to drop, drop rates to spike, and the distributed All-to-All collective barrier (T_barrier = max_i τ_i) to soar, throttling cluster throughput.