Incident Postmortems & Architecture Studies
Real-world failure autopsies, cascading outage analyses, and enterprise architectural studies dissecting GPU memory fragmentation, RL optimization collapse, and throughput bottlenecks with peer-verified remediations.
64x H100 cluster halted · Policy loss NaN · Rollout ceiling 16k tokens
Division by zero in group advantage normalization during uniform rollout failure (std(R) = 0)
Zero-variance clamp guard (std(R) < ε => A=0) · Token generation budget caps · Format penalty
34,210 requests stalled · 102.3x P99 TPOT spike (18.4s) · 1,200 concurrent sessions
Non-contiguous block tail internal fragmentation combined with default host swap triggering PCIe bus thrashing
Zero host swap (--swap-space 0) · Chunked prefill co-scheduling · Radix prefix caching