System breakdown

Inside Local LLM Inference: GGUF Quantization, FlashInfer Kernels & RadixTree KV Reuse

An evidence-audited, 20-diagram interactive system breakdown tracing GGUF super-block quantization (Q4_K_M, IQ4_XS), SIMD & FlashInfer CUDA kernel dequantization, SGLang RadixTree prefix caching, speculative decoding verification loops, and NUMA-aware CPU/GPU memory offloading.

28 min read Verified 2026-07-21 4 primary sources
Canonical System Breakdown
Local LLM Inference Engines (Ollama, llama.cpp & SGLang) Architecture Evidence
RELEASE b2850
5 CHAPTERS · 15 NODES · 10 EDGES
Ingress Security
HMAC SHA256
Constant-time verify (<50ms)
Async Queueing
Sidekiq + Redis
Multi-queue priority isolation
Real-time Fanout
ActionCable Pub/Sub
Redis backplane state broadcast
AI Action Engine
ActionService LLM
Streaming copilot replies (SSE)
Core Stack:Rails 7.1Sidekiq 7Redis 7.2PostgreSQL 16Vue.js 3
5 Verified Architectural Claims
Archify Interactive Mapv2.16.0

Inside Local LLM Inference: GGUF Quantization, FlashInfer Kernels & RadixTree KV Reuse — Archify Interactive System Map

Hotkeys:R Route ProbeL Role LensP Play StoryS Visual StyleT ThemeF Stage