vLLM PagedAttention, KV Cache Block Allocator & Engine Architecture

vLLM PagedAttention, KV Cache Block Allocator & Engine Architecture An architecture diagram generated by Archify. API Client · OpenAI HTTP / SSE · Architecture component · v1/entry API Client OpenAI HTTP / SSE v1/entry AsyncLLM Engine · Request Router · vLLM Host Engine Control Plane · v1/engine AsyncLLM Engine Request Router v1/engine Chunked Scheduler · Prefill & Decode · vLLM Host Engine Control Plane · Co-Scheduling Chunked Scheduler Prefill & Decode Co-Scheduling KVCacheManager · Virtual Page Table · vLLM Host Engine Control Plane · Block Allocator KVCacheManager Virtual Page Table Block Allocator GPU Model Runner · Worker Execution · GPU Acceleration Plane · Worker Loop GPU Model Runner Worker Execution Worker Loop CUDA Graph Runner · Static Graph Replay · GPU Acceleration Plane · Zero Overhead CUDA Graph Runner Static Graph Replay Zero Overhead PagedAttention · CUDA Kernel · GPU Acceleration Plane · VRAM Lookup PagedAttention CUDA Kernel VRAM Lookup Physical GPU HBM · Non-Contiguous Pages · GPU Acceleration Plane · HBM3 / GDDR6 Physical GPU HBM Non-Contiguous Pages HBM3 / GDDR6 Shared GQA Cache · Grouped-Query Attn · GPU Acceleration Plane · 8x Bandwidth Shared GQA Cache Grouped-Query Attn 8x Bandwidth requests enqueue alloc blocks batch replay graph dispatch fetch KV gqa reduce vLLM Host Engine Control Plane GPU Acceleration Plane Legend Backend Database Cloud External

Non-Contiguous VRAM Paging

  • • Partitions GPU memory into 16-token physical blocks
  • • Eliminates 96% internal and external memory fragmentation

Chunked Prefill Co-Scheduling

  • • Slices long prompt prefills into discrete token chunks
  • • Co-schedules prefill chunks with decode passes to saturate tensor cores

Static CUDA Graph Replay

  • • Pre-captures decode forward pass kernel sequences into GPU execution graphs
  • • Reduces CPU launch latency to near zero