FullStack Publications & Research Hub
Peer-Reviewed Evidence Anchors
Publications & Research Library
Canonical academic paper teardowns, real-world incident failure postmortems, production architecture reference guides, and strategic market maps with primary evidence citations.
All Publications84
Seminal Papers11
Incident Postmortems8
Architecture Guides30
System Teardowns25
Research Paper
16 min readDefinitive retrieval-systems teardown of ColBERTv2 (Santhanam et al., arXiv:2112.01488): bridging single-vector bi-encoder compression loss and cross-encoder latency via token-level MaxSim late interaction, centroid-residual b-bit quantization, and PLAID inverted-list pruning.
29 Sep 20262 sources
Read ArticleResearch Paper
18 min readDefinitive mathematical and distributed-systems teardown of DeepSeekMath (arXiv:2402.03300): eliminating the PPO value critic via Group Relative Policy Optimization (GRPO), intra-group advantage normalization, unbiased per-token KL divergence estimation, and outcome vs. process supervision.
29 Sep 20262 sources
Read ArticleResearch Paper
17 min readDefinitive mathematical and systems teardown of QLoRA (Dettmers et al., arXiv:2305.14314): information-theoretically optimal 4-bit NormalFloat (NF4) quantile binning, blockwise k-bit quantization, FP8 Double Quantization of scale constants, and CUDA Unified Memory Paged Optimizers.
29 Sep 20262 sources
Read ArticleResearch Paper
16 min readDefinitive distributed-systems and arithmetic-intensity teardown of Ring Attention (Liu et al., arXiv:2310.01889): overcoming single-GPU HBM limits via ring-topology P2P KV block circulation, overlapped communication and FlashAttention compute, and exact online softmax rescaling.
29 Sep 20262 sources
Read ArticleResearch Paper
17 min readDefinitive mathematical and kernel-level teardown of RoFormer (Su et al., arXiv:2104.09864): encoding relative token distances via multiplicative 2D complex-plane rotations, sparse element-wise slice kernels, long-term decay bounds, and NTK-aware / YaRN context extension.
29 Sep 20262 sources
Read ArticleResearch Paper
16 min readDefinitive mathematical and systems teardown of Speculative Decoding (Leviathan et al., arXiv:2211.17192): overcoming the memory-bandwidth roofline via small draft speculation, parallel target verification, lossless modified rejection sampling, and residual distribution renormalization.
29 Sep 20262 sources
Read ArticleResearch Paper
18 min readTechnical paper teardown of Attention Is All You Need (Vaswani et al., NIPS 2017) detailing Scaled Dot-Product Attention math, Multi-Head Attention projections, Sinusoidal Positional Encoding, and KV cache memory bounds.
29 Jul 20261 source
Read ArticleResearch Paper
16 min readDeepSeek-V3 & R1 Paper Breakdown: Multi-Head Latent Attention, Auxiliary-Loss-Free MoE, and DualPipe
Definitive technical paper breakdown of DeepSeek-V3 and R1 detailing Multi-Head Latent Attention (MLA) low-rank KV compression, auxiliary-loss-free MoE load balancing, and DualPipe pipeline parallelism.
26 Jul 20262 sources
Read ArticleResearch Paper
15 min readDefinitive paper teardown of FlashAttention-3 detailing producer-consumer warp specialization, asynchronous TMA memory loads, FP8 GEMM MMA execution, and inter-warp communication on Hopper GPUs.
26 Jul 20261 source
Read ArticleResearch Paper
14 min readDefinitive paper teardown of vLLM's PagedAttention architecture detailing virtual memory block translation, dynamic copy-on-write sequence forks, and prefix caching.
26 Jul 20261 source
Read ArticleResearch Paper
18 min readPredictive processing offers a useful account of hierarchical inference and error correction, but it is not a shortcut from brain metaphor to system architecture.
16 Jul 20261 source
Read Article