FullStack Publications & Research Hub
Peer-Reviewed Evidence Anchors

Publications & Research Library

Canonical academic paper teardowns, real-world incident failure postmortems, production architecture reference guides, and strategic market maps with primary evidence citations.

All Publications84
Seminal Papers11
Incident Postmortems8
Architecture Guides30
System Teardowns25
Research Paper
16 min read
arXiv:2112.01488
Definitive retrieval-systems teardown of ColBERTv2 (Santhanam et al., arXiv:2112.01488): bridging single-vector bi-encoder compression loss and cross-encoder latency via token-level MaxSim late interaction, centroid-residual b-bit quantization, and PLAID inverted-list pruning.
29 Sep 20262 sources
Read Article
Research Paper
18 min read
arXiv:2402.03300
Definitive mathematical and distributed-systems teardown of DeepSeekMath (arXiv:2402.03300): eliminating the PPO value critic via Group Relative Policy Optimization (GRPO), intra-group advantage normalization, unbiased per-token KL divergence estimation, and outcome vs. process supervision.
29 Sep 20262 sources
Read Article
Research Paper
17 min read
arXiv:2305.14314
Definitive mathematical and systems teardown of QLoRA (Dettmers et al., arXiv:2305.14314): information-theoretically optimal 4-bit NormalFloat (NF4) quantile binning, blockwise k-bit quantization, FP8 Double Quantization of scale constants, and CUDA Unified Memory Paged Optimizers.
29 Sep 20262 sources
Read Article
Research Paper
16 min read
arXiv:2310.01889
Definitive distributed-systems and arithmetic-intensity teardown of Ring Attention (Liu et al., arXiv:2310.01889): overcoming single-GPU HBM limits via ring-topology P2P KV block circulation, overlapped communication and FlashAttention compute, and exact online softmax rescaling.
29 Sep 20262 sources
Read Article
Research Paper
17 min read
arXiv:2104.09864
Definitive mathematical and kernel-level teardown of RoFormer (Su et al., arXiv:2104.09864): encoding relative token distances via multiplicative 2D complex-plane rotations, sparse element-wise slice kernels, long-term decay bounds, and NTK-aware / YaRN context extension.
29 Sep 20262 sources
Read Article
Research Paper
16 min read
arXiv:2211.17192
Definitive mathematical and systems teardown of Speculative Decoding (Leviathan et al., arXiv:2211.17192): overcoming the memory-bandwidth roofline via small draft speculation, parallel target verification, lossless modified rejection sampling, and residual distribution renormalization.
29 Sep 20262 sources
Read Article
Research Paper
18 min read
arXiv:1706.03762
Technical paper teardown of Attention Is All You Need (Vaswani et al., NIPS 2017) detailing Scaled Dot-Product Attention math, Multi-Head Attention projections, Sinusoidal Positional Encoding, and KV cache memory bounds.
29 Jul 20261 source
Read Article
Research Paper
16 min read
arXiv:2412.19437
Definitive technical paper breakdown of DeepSeek-V3 and R1 detailing Multi-Head Latent Attention (MLA) low-rank KV compression, auxiliary-loss-free MoE load balancing, and DualPipe pipeline parallelism.
26 Jul 20262 sources
Read Article
Research Paper
15 min read
arXiv:2407.08608
Definitive paper teardown of FlashAttention-3 detailing producer-consumer warp specialization, asynchronous TMA memory loads, FP8 GEMM MMA execution, and inter-warp communication on Hopper GPUs.
26 Jul 20261 source
Read Article
Research Paper
14 min read
arXiv:2309.06180
Definitive paper teardown of vLLM's PagedAttention architecture detailing virtual memory block translation, dynamic copy-on-write sequence forks, and prefix caching.
26 Jul 20261 source
Read Article
Research Paper
18 min read
doi:10.1098/rstb.2008.0300
Predictive processing offers a useful account of hierarchical inference and error correction, but it is not a shortcut from brain metaphor to system architecture.
16 Jul 20261 source
Read Article