Back to Visual Labs
Lab 19 · Distributed Systems & Hardware Simulator
DeepSeek-V3 & Mixtral 8x7B Architecture

MoE Dynamic Routing & Expert Parallelism Simulator

Explore how distributed Mixture-of-Experts clusters route tokens dynamically across 16 to 64 fine-grained experts hosted on 8 GPU ranks. Observe Top-2 gating softmax, simulate buffer capacity factor limits ($C$), compare traditional auxiliary loss with DeepSeek-V3 auxiliary-loss-free dynamic bias, and inject live chaos to witness distributed All-to-All straggler barrier bottlenecks in real time.

Fine-Grained Experts (E)
16 Experts / 8 GPUs

Mixtral (16E) vs DeepSeek-V3 fine-grained clustering (64E).

Capacity Factor (C)
1.25× Threshold
1.25x

Capacity_E = ⌈(T · K / E) · C⌉. Below 1.25x drops tokens during hotspots.

Balancing Regime
Aux Loss

Mathematical regularization technique to avoid single-expert bottleneck.

Live Chaos Mode:
Aux Balancing Loss
OPTIMAL
0.00128

L_balance = α · K · ∑ f_i · P_i (penalizes gate collapse)

Routing Entropy
UNIFORM
99.9%

4.00 bits (max: uniform 100%)

Token Drop Rate
BUFFER OVERFLOW
21.9%

14 / 64 slots dropped

Straggler Barrier
p99 = 5.45 ms
5.5 ms

All-to-All synchronization wait time (max_i τ_i)

Buffer Slots & Padding Waste
80 Slots Total
Active Compute Slots:50
Wasted Idle Padding:30
Utilization: 62.5%Quota: 5 slots/expert
Gini Load Inequality
G = 0.365
Load Skew Index:Moderate Skew
0.0 (Perfect Equality)1.0 (Total Collapse)
Cluster Throughput
9,174 tok/s
Latency p50 / p95 / p99:5.21 / 5.45 / 5.45 ms
p50 Basep95 JitterBarrier Straggler

Distributed Expert Parallelism Topology (8 GPU Ranks)

All-to-All collective dispatch routes tokens across 8 physical GPU devices. Click any expert or GPU node to inject chaos.

Online Hotspot (+15.0) Offline
G0
GPU Rank 0
Load: 8/10 slots
-2 drop
Expert-00
Code / AST
ONLINE
Queue:3 / 5 tokens
Bias: 0.003.1 ms
::std::sync::Arc[0, 1, 2, ...tail]-> Result<T, E>
Expert-01
Linear Algebra / Softmax
ONLINE
Queue:5 / 5 tokens
Bias: 0.005.3 ms
class TransformerBlocktensor.contiguous()optimizer.step()class TransformerBlock+1 more
G1
GPU Rank 1
Load: 5/10 slots
-2 drop
Expert-02
Kernel / CUDA Mem
ONLINE
Queue:5 / 5 tokens
Bias: 0.005.3 ms
class TransformerBlocktensor.contiguous()optimizer.step()class TransformerBlock+1 more
Expert-03
Reasoning / Knowledge
ONLINE
Queue:0 / 5 tokens
Bias: 0.002.1 ms
G2
GPU Rank 2
Load: 5/10 slots
-1 drop
Expert-04
Syntax / Formatting
ONLINE
Queue:0 / 5 tokens
Bias: 0.002.1 ms
Expert-05
Attention Projections
ONLINE
Queue:5 / 5 tokens
Bias: 0.005.2 ms
∇_θ L_total∫_{-∞}^∞ e^{-x²} dx∫_{-∞}^∞ e^{-x²} dxeigen(W_gate)+1 more
G3
GPU Rank 3
Load: 5/10 slots
-1 drop
Expert-06
FFN Gated Linear
ONLINE
Queue:5 / 5 tokens
Bias: 0.005.2 ms
∇_θ L_total∫_{-∞}^∞ e^{-x²} dx∫_{-∞}^∞ e^{-x²} dxeigen(W_gate)+1 more
Expert-07
Embeddings & Normalization
ONLINE
Queue:0 / 5 tokens
Bias: 0.001.9 ms
G4
GPU Rank 4
Load: 5/10 slots
-2 drop
Expert-08
Code / AST
ONLINE
Queue:0 / 5 tokens
Bias: 0.002.0 ms
Expert-09
Linear Algebra / Softmax
ONLINE
Queue:5 / 5 tokens
Bias: 0.005.5 ms
cudaMemcpyAsyncepoll_wait(ring_fd)posix_memalignSRAM_TILE_LOAD_256K+1 more
G5
GPU Rank 5
Load: 5/10 slots
-2 drop
Expert-10
Kernel / CUDA Mem
ONLINE
Queue:5 / 5 tokens
Bias: 0.005.2 ms
cudaMemcpyAsyncepoll_wait(ring_fd)posix_memalignSRAM_TILE_LOAD_256K+1 more
Expert-11
Reasoning / Knowledge
ONLINE
Queue:0 / 5 tokens
Bias: 0.002.3 ms
G6
GPU Rank 6
Load: 9/10 slots
-4 drop
Expert-12
Syntax / Formatting
ONLINE
Queue:4 / 5 tokens
Bias: 0.005.1 ms
curriculum distillationself-reflective critiquemulti-hop synthesisself-reflective critique
Expert-13
Attention Projections
ONLINE
Queue:5 / 5 tokens
Bias: 0.005.3 ms
curriculum distillationself-reflective critiquein-context few-shotin-context few-shot+1 more
G7
GPU Rank 7
Load: 8/10 slots
Expert-14
FFN Gated Linear
ONLINE
Queue:5 / 5 tokens
Bias: 0.005.4 ms
in-context few-shotin-context few-shotin-context few-shotcurriculum distillation+1 more
Expert-15
Embeddings & Normalization
ONLINE
Queue:3 / 5 tokens
Bias: 0.003.2 ms
::std::sync::Arc[0, 1, 2, ...tail]-> Result<T, E>

Microbatch Token Dispatch Matrix (Top-2 Gating Softmax)

Per-token routing decisions, normalized gating weights (g_1 + g_2 = 1.0), and expert buffer drop statuses.

#Token Vector & TextCategoryTop-1 Expert & Weight (g₁)Top-2 Expert & Weight (g₂)Decision EntropyDispatch Status
1class TransformerBlock
code
Expert-02(G1)50.8%
Expert-01(G0)49.2%
95%
ROUTED (2/2)
2curriculum distillation
natural
Expert-13(G6)53.4%
Expert-12(G6)46.6%
95%
ROUTED (2/2)
3cudaMemcpyAsync
system
Expert-09(G4)50.8%
Expert-10(G5)49.2%
94%
ROUTED (2/2)
4self-reflective critique
natural
Expert-13(G6)53.3%
Expert-12(G6)46.7%
95%
ROUTED (2/2)
5tensor.contiguous()
code
Expert-02(G1)50.8%
Expert-01(G0)49.2%
96%
ROUTED (2/2)
6in-context few-shot
natural
Expert-13(G6)52.4%
Expert-14(G7)47.6%
95%
ROUTED (2/2)
7::std::sync::Arc
punctuation
Expert-00(G0)50.2%
Expert-15(G7)49.8%
98%
ROUTED (2/2)
8in-context few-shot
natural
Expert-13(G6)53.0%
Expert-14(G7)47.0%
95%
ROUTED (2/2)
9[0, 1, 2, ...tail]
punctuation
Expert-15(G7)50.7%
Expert-00(G0)49.3%
98%
ROUTED (2/2)
10epoll_wait(ring_fd)
system
Expert-10(G5)51.2%
Expert-09(G4)48.8%
94%
ROUTED (2/2)
11∇_θ L_total
math
Expert-05(G2)51.2%
Expert-06(G3)48.8%
94%
ROUTED (2/2)
12posix_memalign
system
Expert-10(G5)50.7%
Expert-09(G4)49.3%
95%
ROUTED (2/2)
13optimizer.step()
code
Expert-02(G1)50.2%
Expert-01(G0)49.8%
96%
ROUTED (2/2)
14SRAM_TILE_LOAD_256K
system
Expert-09(G4)50.7%
Expert-10(G5)49.3%
94%
ROUTED (2/2)
15class TransformerBlock
code
Expert-02(G1)50.5%
Expert-01(G0)49.5%
95%
ROUTED (2/2)
16-> Result<T, E>
punctuation
Expert-15(G7)51.4%
Expert-00(G0)48.6%
98%
ROUTED (2/2)
17∫_{-∞}^∞ e^{-x²} dx
math
Expert-05(G2)50.5%
Expert-06(G3)49.5%
94%
ROUTED (2/2)
18in-context few-shot
natural
Expert-13(G6)53.2%
Expert-14(G7)46.8%
95%
ROUTED (2/2)
19∫_{-∞}^∞ e^{-x²} dx
math
Expert-05(G2)50.0%
Expert-06(G3)50.0%
93%
ROUTED (2/2)
20optimizer.step()
code
Expert-02(G1)50.3%
Expert-01(G0)49.7%
95%
ROUTED (2/2)
21posix_memalign
system
Expert-10(G5)50.8%
Expert-09(G4)49.2%
94%
ROUTED (2/2)
22eigen(W_gate)
math
Expert-05(G2)50.7%
Expert-06(G3)49.3%
94%
ROUTED (2/2)
23curriculum distillation
natural
Expert-13(G6)52.9%[DROP]
Expert-14(G7)47.1%
95%
CAPACITY EXCEEDED
24self-reflective critique
natural
Expert-13(G6)53.4%[DROP]
Expert-14(G7)46.6%
95%
CAPACITY EXCEEDED
25multi-hop synthesis
natural
Expert-13(G6)53.4%[DROP]
Expert-12(G6)46.6%
95%
CAPACITY EXCEEDED
26eigen(W_gate)
math
Expert-05(G2)50.4%
Expert-06(G3)49.6%
94%
ROUTED (2/2)
27async function eval
code
Expert-01(G0)51.1%[DROP]
Expert-02(G1)48.9%[DROP]
95%
CAPACITY EXCEEDED
28tensor.contiguous()
code
Expert-02(G1)51.1%[DROP]
Expert-01(G0)48.9%[DROP]
96%
CAPACITY EXCEEDED
29∫_{-∞}^∞ e^{-x²} dx
math
Expert-05(G2)50.5%[DROP]
Expert-06(G3)49.5%[DROP]
94%
CAPACITY EXCEEDED
30self-reflective critique
natural
Expert-13(G6)53.3%[DROP]
Expert-12(G6)46.7%
95%
CAPACITY EXCEEDED
31cudaMemcpyAsync
system
Expert-09(G4)50.6%[DROP]
Expert-10(G5)49.4%[DROP]
94%
CAPACITY EXCEEDED
32All-to-All Dispatch
system
Expert-09(G4)50.7%[DROP]
Expert-10(G5)49.3%[DROP]
94%
CAPACITY EXCEEDED

MoE Router & Distributed Collective Event Stream

Latest 3 dispatch events
7:15:09 PM[tokens-dropped]Microbatch #2: 14 token assignments dropped due to buffer saturation / offline nodes.
7:15:08 PM[preset-loaded]Initialized MoE cluster preset "mixtral-8x7b-balanced" with 16 experts (C=1.25x, auxiliary-loss).
7:15:08 PM[tokens-dropped]Microbatch #1: 22 token assignments dropped due to buffer saturation / offline nodes.
1

Auxiliary Loss vs DeepSeek-V3 Dynamic Bias

Switch between Aux Loss (L_balance = α · K · ∑ f_i · P_i) and DeepSeek Bias (b_i ← b_i + γ(1/E - f_i)). Observe how DeepSeek-V3 autonomously steers traffic away from congested experts without injecting surrogate gradient penalties into the core representation loss.

Verifies Loss-Free Convergence
2

Capacity Factor (C) & Buffer Waste

Modulate the Capacity Factor slider from 1.0× to 2.0×. At C=1.0× (zero headroom), even slight Poisson variance causes token drops. At C=2.0×, drops fall to 0%, but idle padding slots consume up to 50% of allocated GPU tensor buffers.

Verifies Memory/Drop Trade-Off
3

Celebrity Hotspot Skew & Straggler Barrier

Click Celebrity Hotspot (+15.0). Watch tokens collapse into hot experts, causing entropy to drop, drop rates to spike, and the distributed All-to-All collective barrier (T_barrier = max_i τ_i) to soar, throttling cluster throughput.

Verifies All-to-All Straggler Bottleneck