paper breakdown

RoFormer & RoPE Paper Breakdown: Rotary Position Embeddings, Complex-Plane Inner Products, and Long-Context Scaling

Definitive mathematical and kernel-level teardown of RoFormer (Su et al., arXiv:2104.09864): encoding relative token distances via multiplicative 2D complex-plane rotations, sparse element-wise slice kernels, long-term decay bounds, and NTK-aware / YaRN context extension.

17 min readVerified 2026-09-292 primary sources
Technical publication illustration.

RoFormer & Rotary Position Embeddings (RoPE) Paper Breakdown

A mathematical and kernel-level breakdown of RoFormer: Enhanced Transformer with Rotary Position Embedding (Su, Lu, Pan, Murtadha, Wen, & Liu, arXiv:2104.09864), the foundational paper behind the positional encoding used in virtually every modern frontier LLM (LLaMA 1/2/3/4, Qwen, Mistral, Gemma, and DeepSeek-V3/R1).


1. The Fundamental Unification Problem: Absolute vs. Relative Positions

In self-attention, given token embeddings x_m at position m and x_n at position n, we apply position-aware query and key transformations q_m = f_q(x_m, m) and k_n = f_k(x_n, n).

Additive absolute position encodings (f(x_m, m) = x_m + p_m) pollute the semantic representation with position vectors and fail to preserve translation invariance. Conversely, explicit relative position biases (q_m^T k_n + b_{m-n}) require materializing an O(N^2) position matrix, breaking linear attention and complicating KV cache indexing.

Su et al. pose the foundational functional equation of RoPE: Find an absolute point-wise operation f_q(x_m, m) and f_k(x_n, n) whose inner product depends strictly on the embeddings x_m, x_n and their relative distance m - n:

text(2 lines)
1< f_q(x_m, m), f_k(x_n, n) > = g( x_m, x_n, m - n )

2. Complex-Plane Derivation & 2D Orthogonal Rotation

Consider a 2-dimensional query/key vector q = (q_0, q_1) identified as a complex number q_0 + i * q_1 = ||q|| * exp(i * phi_q). In the complex plane, multiplying a vector by Euler's complex exponential exp(i * m * theta) rotates its phase angle by m * theta without altering its Euclidean norm:

text(3 lines)
1f_q(x_m, m) = (W_q * x_m) * exp(i * m * theta)
2f_k(x_n, n) = (W_k * x_n) * exp(i * n * theta)

Taking the real part of the Hermitian inner product Re[ f_q(x_m, m) * conj(f_k(x_n, n)) ]:

text(3 lines)
1Re[ q_m * conj(k_n) * exp(i * m * theta) * exp(-i * n * theta) ]
2 = Re[ (q_m * conj(k_n)) * exp(i * (m - n) * theta) ]

In real R^2 coordinates, multiplying by exp(i * m * theta) = cos(m * theta) + i * sin(m * theta) is isomorphic to left-multiplying by the 2 x 2 planar rotation matrix:

text(3 lines)
1R(m * theta) = [ cos(m * theta) -sin(m * theta) ]
2 [ sin(m * theta) cos(m * theta) ]

Because rotation matrices satisfy R(a)^T * R(b) = R(-a) * R(b) = R(b - a), the attention logit between position m and position n simplifies algebraically to:

text(3 lines)
1( R(m * theta) * q )^T * ( R(n * theta) * k ) = q^T * R(m * theta)^T * R(n * theta) * k
2 = q^T * R( (n - m) * theta ) * k

The absolute token indices m and n vanish completely from the inner product, leaving only the relative token offset n - m!


3. Generalization to d Dimensions & O(d) Slice Kernel

For a head dimension d (where d is even, e.g., d = 128), RoPE partitions the d-dimensional vector into d / 2 independent 2D subspaces (2j, 2j + 1) for j in {0, 1, ..., d/2 - 1}, constructing a block-diagonal orthogonal matrix R_{Theta, m}^d:

text(2 lines)
1R_{Theta, m}^d = diag( R(m * theta_0), R(m * theta_1), ..., R(m * theta_{d/2 - 1}) )

Where the geometric frequency spectrum Theta follows the sinusoidal progression with base b = 10000.0 (or 500000.0 in LLaMA-3):

text(2 lines)
1theta_j = b^( -2 * j / d ) = 1.0 / ( 10000.0 ^ (2 * j / d) )

Zero-GEMM Element-Wise Implementation

Never multiply by the dense d x d matrix R_{Theta, m}^d (O(d^2) FLOPs). Because R_{Theta, m}^d has only two non-zero elements per row, RoPE is executed in O(d) register operations via the rotate_half element-wise decomposition:

text(3 lines)
1R_{Theta, m}^d * x = [ x_0, x_1, x_2, x_3, ..., x_{d-2}, x_{d-1} ] (*) [ cos(m*theta_0), cos(m*theta_0), ..., cos(m*theta_{d/2-1}), cos(m*theta_{d/2-1}) ]
2 + [-x_1, x_0,-x_3, x_2, ..., -x_{d-1}, x_{d-2} ] (*) [ sin(m*theta_0), sin(m*theta_0), ..., sin(m*theta_{d/2-1}), sin(m*theta_{d/2-1}) ]
python(22 lines)
1import math
2
3def apply_rope_2d_slice(
4 vec: list[float], # Length d (even)
5 pos: int, # Token index m
6 base: float = 10000.0 # RoPE theta base
7) -> list[float]:
8 d = len(vec)
9 out = [0.0] * d
10 for j in range(d // 2):
11 theta_j = 1.0 / (base ** (2.0 * j / d))
12 angle = pos * theta_j
13 cos_a = math.cos(angle)
14 sin_a = math.sin(angle)
15
16 x0 = vec[2 * j]
17 x1 = vec[2 * j + 1]
18 # 2D planar rotation: [x0*cos - x1*sin, x0*sin + x1*cos]
19 out[2 * j] = x0 * cos_a - x1 * sin_a
20 out[2 * j + 1] = x0 * sin_a + x1 * cos_a
21 return out

4. Relative Distance Decay & Long-Context Scaling (PI, NTK, YaRN)

Using an Abel transformation (summation by parts) across the d / 2 frequencies, Su et al. prove that the upper bound of the inner product |q_m^T k_n| decays naturally as the relative distance |m - n| increases—imparting an inductive recency bias while preserving sharp local attention.

Why Context Extrapolation Fails & How NTK / YaRN Fix It

Each subspace j has a wavelength lambda_j = 2 * pi / theta_j = 2 * pi * b^(2j / d):

  • High-frequency subspaces (j -> 0): Wavelength lambda_0 = 2 * pi ≈ 6.28 tokens. Encodes local syntax and adjacent word order.
  • Low-frequency subspaces (j -> d/2): Wavelength lambda_{d/2} = 2 * pi * 10000 ≈ 62,831 tokens. Encodes global document position.

| Extension Method | Mathematical Transformation on theta_j or m | Impact on High vs. Low Frequencies | |---|---|---| | Position Interpolation (PI) | Scales position m' = m / s (s = L_ext / L_train) uniformly across all j | Compresses local high-frequency resolution by s; requires ~1,000 steps of fine-tuning | | NTK-Aware RoPE Scaling | Scales base b' = b * s^(d / (d - 2)) instead of scaling m linearly | Interpolates low frequencies while extrapolating high frequencies to preserve local token separation | | YaRN (NTK-by-Parts + Temp) | Ramp function (1 - gamma(r)) * (theta_j / s) + gamma(r) * theta_j + logit scale sqrt(1/t) | Leaves high frequencies (lambda_j < L_train) untouched (gamma=1), interpolates only low frequencies |


5. Production Systems Invariants

  1. KV Cache Position Invariance: Because k_n = R(n * theta) * k embeds position n directly into the stored Key vector in the PagedAttention KV cache, cached keys never need to be recomputed as new tokens m > n are generated—except under sliding-window streaming eviction (StreamingLLM) where evicted tokens require re-rotating keys to contiguous virtual positions.
  2. Interleaved vs. Split-Half Layout: HuggingFace LLaMA permutes W_q and W_k rows so the 2D pairs are stored at (j, j + d/2) (split-half) rather than (2j, 2j+1) (interleaved GPT-J style), enabling coalesced 128-bit SIMD vector loads in CUDA/Triton kernels.