Back to Visual Labs
Lab 19 · Hardware SRAM Tiling Simulator
FlashAttention-3 Hopper H100 Implementation

FlashAttention-3 SRAM Tiling & Online Softmax Lab

Explore how GPU memory hierarchy transforms long-context LLM attention. Simulate Br × Bc SRAM tile streaming, watch online softmax rescaling factor alpha = exp(m_prev - m_curr) eliminate intermediate HBM reads and writes, observe causal tile skipping, and inspect real-time mathematical proof of how FlashAttention turns an O(N²) VRAM crash (137 GB at 32k) into a strictly bounded O(N) linear execution.

Initializing FlashAttention-3 Hardware Simulator Kernel...
1

The O(N²) VRAM Blowout (32k → 128k)

Load the 32k Long-Context preset. Observe how standard attention requires 137.4 GB of HBM just to materialize S and P, triggering an instant OOM crash on an 80 GB GPU. In contrast, FlashAttention-3 uses only 1.07 GB in HBM and fixed 160 KB in SRAM.

Verifies O(N²) vs O(N) Footprint Invariant
2

Online Softmax Rescaling Factor alpha

Switch to the Vector Inspector tab and click Step Next. Watch row statistics update dynamically. Whenever a new tile contains a larger logit m_tilde > m, the engine scales previous accumulators by alpha = exp(m_prev - m_curr), guaranteeing exact softmax without multi-pass HBM roundtrips.

Verifies Monotonic Normalizer Exactness
3

Hopper H100 TMA & Roofline Saturation

Toggle Stepping Granularity to Micro-Op. Trace how Hopper's Asynchronous Tensor Memory Accelerator (TMA) streams K_j, V_j tiles directly into SRAM, allowing Tensor Cores to achieve high operational intensity (> 256 FLOPs/Byte) in the compute-bound regime.

Verifies Compute-Bound Tensor Core Regimes