FlashAttention-3 SRAM Tiling & Online Softmax Lab
Explore how GPU memory hierarchy transforms long-context LLM attention. Simulate Br × Bc SRAM tile streaming, watch online softmax rescaling factor alpha = exp(m_prev - m_curr) eliminate intermediate HBM reads and writes, observe causal tile skipping, and inspect real-time mathematical proof of how FlashAttention turns an O(N²) VRAM crash (137 GB at 32k) into a strictly bounded O(N) linear execution.
The O(N²) VRAM Blowout (32k → 128k)
Load the 32k Long-Context preset. Observe how standard attention requires 137.4 GB of HBM just to materialize S and P, triggering an instant OOM crash on an 80 GB GPU. In contrast, FlashAttention-3 uses only 1.07 GB in HBM and fixed 160 KB in SRAM.
Online Softmax Rescaling Factor alpha
Switch to the Vector Inspector tab and click Step Next. Watch row statistics update dynamically. Whenever a new tile contains a larger logit m_tilde > m, the engine scales previous accumulators by alpha = exp(m_prev - m_curr), guaranteeing exact softmax without multi-pass HBM roundtrips.
Hopper H100 TMA & Roofline Saturation
Toggle Stepping Granularity to Micro-Op. Trace how Hopper's Asynchronous Tensor Memory Accelerator (TMA) streams K_j, V_j tiles directly into SRAM, allowing Tensor Cores to achieve high operational intensity (> 256 FLOPs/Byte) in the compute-bound regime.