Reading Notes — Attention & KV-Cache Mechanics
A running notebook on how recent work decides what a query is allowed to look at — masks, sliding windows, and cache eviction.
- InfoKV — Information-Aware KV Cache Compression for Long Reasoning · arXiv:2606.26875
- Causal-rCM — autoregressive diffusion distillation for streaming video & world models (block-causal attention) · arXiv:2606.25473
- Unlimited OCR Works — Reference Sliding Window Attention (R-SWA) · arXiv:2606.23050
SnapKV (2404.14469) · PyramidKV (2406.02069) · Expected Attention (2510.00636) · StreamingLLM / attention sinks (2309.17453) · R-KV (2505.24133) · ShortGPT (2403.03853) · H2O (2306.14048) · DeepSeek-V2/V3 (MLA) · Arxiv-Summarization (Cohan et al. 2018)
Two of today's papers are anonymous EMNLP submissions, not yet public. They are referenced only generically here; titles, results, and identifying detail are omitted.
The idea in Information-Aware KV Cache Compression for Long Reasoning reminds me of two things I read recently: the block-causal attention used in arXiv:2606.25473 (Causal-rCM), and the R-SWA method from Baidu's Unlimited-OCR. They feel related — could they be combined?
The instinct is right, and the reason it's right is worth stating precisely: all three are answers to the same question — a query cannot afford to attend to everything, so what does it attend to? Each answers at a different level. That's exactly why they compose rather than collide.
First, what each one actually is
| Mechanism | How it decides | Acts on | KV vs. output length | Long-range recall via |
|---|---|---|---|---|
| Info-Aware KV Compression 2606.26875 |
Content-adaptive — information-theoretic scoring (entropy / info-content / redundancy) | The cache — evicts or compresses stored KV | ~O(budget), but which tokens are kept is dynamic | Retains the highest-information past tokens |
| Block-causal attention 2606.25473 |
Structural — fixed by chunk topology | The mask — restricts visibility | O(L) in principle | Full causal history of past chunks |
| R-SWA 2606.23050 |
Structural — sliding window + fixed reference anchor | Both — the window is a mask that pins the cache to O(1) | O(1) | A fixed reference anchor (the compressed source) + recent window |
Caveat on grounding: I'm working from the title and a one-line briefing for 2606.26875 plus its immediate neighbours (entropy-guided caching, information-bottleneck eviction, redundancy-aware R-KV). The "information-theoretic scoring" characterization is the family it belongs to, not a verbatim description of its specific algorithm.
The common thread cuts along two orthogonal axes
Structural (position/topology, fixed in advance) vs. content-adaptive (depends on the data). Block-causal and R-SWA are structural; information-aware compression is content-adaptive.
The mask (which positions are visible, recomputed every forward pass) vs. the cache (which KV entries are actually stored). Block-causal touches only the mask; info-aware touches only the cache; R-SWA touches both — its window is a mask, but because the width is fixed it also bounds the physical cache to O(1).
This is why they're complementary, not redundant: structural sparsity (block-causal, sliding window) pins the cache shape cheaply but blindly; content-aware compression spends compute to keep the right tokens but doesn't by itself guarantee O(1). Combining them isn't circular — each supplies what the other lacks.
How to combine them
Use the information-aware policy to decide what fills R-SWA's fixed budget, and make that decision at block-causal chunk granularity.
The division of labour is clean:
Structural layer — R-SWA-style
Guarantees O(1) cache and cheap attention: a sliding window over recently generated tokens (local coherence) plus a fixed-capacity "reference bank" of anchor slots.
Content layer — Information-Aware
Decides what occupies the reference bank. In OCR the anchor is the compressed source image — and that works because OCR is near-monotonic copying: what you need to look back at lives in the source and is local. Reasoning is non-monotonic: a conclusion 3,000 tokens later may depend on a lemma established at the start. So populate and refresh the bank with the highest-information past tokens by the information-theoretic criterion. This is StreamingLLM's "attention sink + window" generalized: the sink slots are content-selected and refreshed, not a hard-coded "first 4 tokens."
Granularity knob — Block-causal
Don't run selection per token. Generate a chunk, then at the chunk boundary run the information-aware compression to decide which of the chunk's KV (or a single summary token) get promoted into the reference bank; let the window cover the last k chunks. Chunk-level amortization is what keeps the (non-trivial) scoring overhead from eating the O(1) savings — and it matches the streaming structure Causal-rCM already uses.
Net effect: OCR-like O(1) memory with reasoning-grade long-range recall — a constant budget, but spent on the informative tokens rather than merely the recent ones.
Where it could break — the honest caveats
- Task structure decides feasibility. R-SWA works for OCR because a persistent external reference (the source image) carries long-range memory for free. Reasoning has no external source — the thing to look back at is self-generated and growing — so the content layer must manufacture the reference OCR gets for free. Choose wrong and you evict a lemma you'll need, with no ground-truth source to fall back on.
- Don't import block-causal's bidirectional interior. In video, intra-chunk bidirectional attention is fine (frames are co-generated). In autoregressive reasoning a token shouldn't attend to future tokens in its own chunk. Borrow the chunk-as-unit idea but keep the chunk interior causal.
- Trained-in beats bolt-on. R-SWA replaces attention layers and is trained that way; most info-aware compression is a training-free inference-time add-on. Bolting a training-free eviction policy onto a model trained for full attention degrades reasoning (the motivation behind R-KV and "Hold Onto That Thought"). The strongest version trains the model with the block-causal + window + content-bank attention, so it learns to emit promotable summary tokens at chunk boundaries — i.e., to externalize its own working memory into the bank. That lines up with Unlimited-OCR's "human parsing working memory" framing and the KV-distillation / latent-reasoning line (KaVa).
Prior art to anchor it
This isn't built from nothing: MorphKV and LaCache already pursue constant-size caches that keep context-correlated older tokens; R-KV does redundancy-aware selection; StreamingLLM is the degenerate case of R-SWA where the "reference" is just the first few tokens. The proposal here is essentially R-SWA's O(1) container + R-KV/info-aware selection for the anchor slots + block-causal chunk granularity.
2606.25473 — Causal-rCM (frame-wise & chunk-wise / block-causal streaming diffusion)
2606.23050 — Unlimited OCR Works (Reference Sliding Window Attention, R-SWA)
context: StreamingLLM · R-KV (2505.24133) · MorphKV · LaCache · CapKV (information-bottleneck eviction)
Summarize the whole pipeline of Information-Aware KV Cache Compression for Long Reasoning (InfoKV) and draw a flowchart — why it's done, what problems exist, the fixes, the gaps in prior methods. Then a batch of specifics: why is Sᵢ = Aᵢ + softmax(H)ᵢ shaped this way (gradient control? mapping to [0,1]?); what gets back-propagated / which are learnable parameters; is KL the only divergence option; why the coefficient in front of formula (4) — is that L1-normalization; can the cross-layer cosine distance serve as a layer-wise redundancy diagnostic; what is "Arxiv-Summarization"; and would this work better on decoder-only models like ChatGPT — how does ChatGPT handle long reasoning under limited KV cache?
Correction to Entry-prior diagram: an earlier sketch labelled Dᵢ = 1 − cos(hˡ, hᴸ) as the "layer-wise budget." That was wrong. Reading the full paper: D is a factor of the entropy score (Eᵢ = Dᵢ · Hᵢ, eq 7); the actual per-layer budget allocation is a separate, optional adaptive variant (eqs 9–10). The paper has two distinct "layer-wise" ideas, and conflating them is the trap. The corrected pipeline is below.
The logic chain: motivation → method → selection
Motivation — what's wrong with prior work
Attention-based KV compression (SnapKV, PyramidKV, FastKV) scores tokens by how much attention they receive from a recent observation window. The flaw is that this is backward-looking and short-range: important-to-recent-context is not the same as important-to-future-reasoning. In long decoding, the reasoning path keeps evolving, so the mismatch compounds. To demonstrate it, InfoKV defines Forward Influence (eq 3): remove a token from the cache and measure the KL divergence between the original and ablated predictive distributions over a future chunk. The finding (Figs 1–2): attention-selected tokens influence only nearby future and decay fast; entropy-selected tokens exert stronger, more persistent influence on distant future. Hence — combine the two.
Method — InfoKV
The entropy score multiplies two orthogonal axes: along the sequence, token predictive uncertainty Hᵢ (eq 2, computed with top-256 restricted entropy for stability); across depth, representation evolution Dᵢ = 1 − cos(hˡ, hᴸ) (eq 6, with a bias τ so the final layer isn't zeroed). Their product is the per-layer entropy score Eᵢ = Dᵢ · Hᵢ (eq 7). This is fused with the attention score Aᵢ (eq 4, taken from the last layer) via a tunable convex combination Sᵢ = α·Aᵢ + (1−α)·softmax(E)ᵢ (eq 8, default α = 0.9). Each layer keeps its top-S tokens.
The form Sᵢ = Aᵢ + softmax(H)ᵢ (the image from the prior turn) is eq 5 — the preliminary version used only in the motivating analysis, fusing attention with raw entropy. The deployed method is eq 8: it swaps raw H for the entropy score E = D·H and adds the balance weight α.
Formula reference
| Eq | Expression | Role |
|---|---|---|
| 2 | Hᵢ = −Σ p̂ log p̂ (top-k) | token-level informativeness (predictive uncertainty) |
| 3 | I = mean·DKL(p ‖ p−i) | Forward Influence — diagnostic only, not in the pipeline |
| 4 | Aᵢ = (1/|W|)·Σ Attn(qₜ,kᵢ) | mean attention over the observation window (last layer) |
| 5 | Sᵢ = Aᵢ + softmax(H)ᵢ | preliminary fusion (analysis section) |
| 6 | Dᵢ = 1 − cos(hˡ, hᴸ) | cross-layer representation evolution (+ bias τ) |
| 7 | Eᵢ = Dᵢ · Hᵢ | entropy score = informativeness × evolution |
| 8 | Sᵢ = α·Aᵢ + (1−α)·softmax(E)ᵢ | the InfoKV score (α = 0.9) |
| 9–10 | kₗ = (Ē⁽ˡ⁾ / ΣĒ)·B | adaptive per-layer budget — optional, less robust |
The specifics
Why Sᵢ = Aᵢ + softmax(H)ᵢ? Not gradient control, not min-max
This is training-free inference scoring — there are no gradients to explode. The paper states it plainly: softmax is used to normalize the scale of entropy along the sequence dimension before adding it to attention. Entropy H is unbounded (nats); attention A already lives in [0,1] and roughly sums to 1 over keys. The two are on incommensurate scales, so raw addition is meaningless. Softmax maps H to (0,1), summing to 1 over the sequence — making it a distribution comparable to A. So "maps to [0,1]" is morally right, but the precise purpose is scale-matching across the sequence axis. Addition (not multiplication) is an OR-style fusion: keep a token flagged by either signal. Eq 8 generalizes this to a tunable convex blend.
What is learnable / back-propagated? Nothing
InfoKV is a training-free, plug-in method on a frozen model. Every quantity (A, H, D, E, S) is read off the forward pass — last-layer attention weights, predictive-distribution entropy, hidden states. No parameter is updated. The "knobs" are hyperparameters, not learned weights: α = 0.9, τ = 1.0, top-k = 256, observation window = 64, compression every 1024 tokens — all chosen by ablation (Figs 5–6, Appendix B). This is the key fact for judging deployability: it bolts onto inference, it doesn't train anything.
KL divergence — alternatives
KL appears in eq 3 (Forward Influence). Forward KL DKL(p ‖ p′) is the natural choice here because it's an expectation under the original distribution p — exactly "how much does ablating this token perturb the reference." Alternatives: JS divergence (symmetric, bounded), total variation / Hellinger / χ² (other f-divergences), Wasserstein (geometry-sensitive), and cross-entropy / mutual information. Note the paper also implicitly uses a distributional-style distance elsewhere — the cosine distance of eq 6 — but on hidden states rather than probabilities.
The coefficient in eq 4 — averaging, not L1-normalization
1/(r₀ − l₀ + 1) is the reciprocal of the window length, so Aᵢ is the arithmetic mean attention key i receives across the window's queries. Its purpose is window-size invariance and aggregation — not [0,1] scaling, and not L1-normalization. (L1-normalization divides by the sum of the values to make them sum to 1; this divides by the count to take a mean.) A neat side-effect: since each query's attention is a softmax over keys (each row sums to 1), the per-key average also sums to 1 — so Aᵢ happens to be a distribution. But the operation is averaging. The same 1/(r−l+1) recurs in eq 3 for the same reason: a mean over a future chunk.
Cosine distance as a layer-wise redundancy probe — usable, with caveats
Eq 6 is a clean, training-free per-token, per-layer convergence/redundancy signal grounded in the residual stream: if an early-layer representation already aligns with the final layer, the token has largely converged and likely carries little additional information; large shifts signal unresolved semantics. Three caveats before reusing it: (1) cosine ignores magnitude, yet residual norm grows with depth — pair it with an update norm ‖hˡ − hˡ⁻¹‖ or a logit-lens / tuned-lens KL. (2) "early convergence ⇒ low future info" is heuristic — some tokens converge early because they're trivial (function words). That's exactly why the paper never uses D alone but as D·H — convergence multiplied by informativeness. (3) The signal is architecture-dependent (see next point). Lesson for redundancy-diagnosis work: don't use a convergence signal in isolation; gate it with an informativeness signal.
"Early/middle layers richer, higher layers redundant" as a pruning prior — real but architecture-bound
This is §5.2, and it motivates the adaptive budget (eq 10: allocate KV budget proportional to a layer's accumulated entropy). But the paper itself walks it back: adaptive allocation helped R1-Distill-Llama-8B yet hurt R1-Distill-Qwen-7B (Table 2) — over-imbalanced budgets over-compress some layers and destabilize long-range reasoning. So they keep uniform as default and flag architecture-aware allocation as future work. For layer-wise pruning: the directional prior holds, but its strength is model-specific — measure each target model's per-layer entropy curve rather than treating it as a universal constant.
What is Arxiv-Summarization?
The long-document summarization dataset from Cohan et al., 2018 (NAACL, A Discourse-Aware Attention Model for Abstractive Summarization of Long Documents): inputs are full arXiv papers, targets are their abstracts. InfoKV uses it only for the Forward-Influence diagnostic (100 documents, Figs 1–2) — it supplies long documents to probe forward influence — not as a main benchmark.
The benchmark set used here
Long prefilling: LongReason (Ling et al., 2025; synthetic long-context reasoning via context expansion) on Llama-3.1-8B and Llama-3.2-3B, at 16k/32k/64k, cache 40%/20%, with and without CoT; baselines SnapKV, PyramidKV, Expected Attention. Long decoding: IFEval (instruction following), AIME 2024 (math), LiveCodeBench (code) on DeepSeek-R1-Distill-Qwen-7B and -Llama-8B, max output 32,768, compression every 1024 tokens, cache 25%/12.5%; baseline RPC.
Decoder-only models / ChatGPT
A framing correction first: the tested models (Llama-3.1/3.2, DeepSeek-R1-Distill) are decoder-only, and so are GPT / o-series. KV-cache compression only makes sense for autoregressive causal decoding, so InfoKV is already a decoder-only method — there's no encoder-decoder comparison to be "better" than. The real transfer question (to a model like GPT-5/o-series) has three constraints: (1) it needs white-box access to attention weights, hidden states, and the predictive distribution — so it cannot be bolted onto a closed API from outside; only the provider could apply it internally. (2) Modern models use GQA/MLA, which changes KV structure and may require adapting the per-layer/per-head budgeting. (3) Transfer is non-uniform — the paper's own adaptive variant helped one distill and hurt another.
As for how ChatGPT handles long reasoning under a limited cache: the serving internals are not publicly documented, so the specific compression algorithm is unknown. Generally, production long-context serving uses PagedAttention-style memory management, prefix/prompt caching, KV quantization, attention-sink / sliding-window streaming, and hard context caps; for reasoning models the chain-of-thought tokens occupy the KV during a turn (billed, usually hidden) and aren't retained across turns. The operative point remains: InfoKV-style methods can't be applied to ChatGPT externally because the weights and intermediate tensors aren't exposed.
context: SnapKV · PyramidKV · FastKV · Expected Attention · RPC · LongReason (Ling 2025) · Arxiv-Summarization (Cohan 2018) · SH2 (Kai 2024) · SeLaR (Fu & Luo 2026)
Explain carefully what "along the sequence dimension" means and what "scale alignment along the sequence dimension" is; and walk through why Aᵢ sums to 1 along the key dimension (the averaged-attention-stays-a-distribution argument). Did I mean only decoder-only models can use these methods — not encoder-decoder? Make a comparison table of SnapKV / PyramidKV / Expected Attention. What is MLA — is it the same as GQA — and add MHA to a comparison table. And explain in detail each of: PagedAttention, prefix/prompt caching, KV quantization, attention sink / sliding window, context-window limits.
Tensor dimensions, and why softmax(H) but not softmax(A)
A sequence of hidden states is a matrix of shape [T, d]. The sequence dimension is T — indexing over token positions (0 … T−1). The feature/hidden dimension is d — the channels inside a single token. Which axis an operation runs "along" determines what it aggregates over: along the sequence (attention's softmax over keys; InfoKV's softmax(H)), along the features (LayerNorm; or a coefficient of variation CV(|hᵢ|) computed over a token's d entries), or along depth (InfoKV's Dᵢ = 1 − cos(hˡ, hᴸ)).
"Scale alignment along the sequence dimension": H is a length-T vector (one entropy per token). Applying softmax along the sequence axis exponentiates each token's entropy and divides by the sum over all T tokens, yielding weights that sum to 1 across the token axis — putting H on the same ruler as the attention weights, which are also a distribution over the sequence. Only then is adding them meaningful.
Why Aᵢ already sums to 1. For a fixed query t, attention is a softmax over keys, so Σᵢ Attn(qₜ, kᵢ) = 1. With Aᵢ = (1/W)·Σₜ Attn(qₜ, kᵢ) (W = window size), summing over keys and swapping the order of summation gives Σᵢ Aᵢ = (1/W)·Σₜ (Σᵢ Attn(qₜ,kᵢ)) = (1/W)·Σₜ 1 = (1/W)·W = 1. Intuition: each query spends a total attention budget of 1 across keys; Aᵢ is the average share key i receives, and an average of distributions is still a distribution. Tiny example — query1 = [0.7, 0.2, 0.1], query2 = [0.1, 0.6, 0.3] → average = [0.4, 0.4, 0.2], which sums to 1. This is exactly why A needs no softmax (it's already a sequence-normalized distribution) while raw H (unbounded nats) does.
Decoder-only vs encoder-decoder — the correction
The dependence is on autoregressive (causal) decoding, not on the literal "decoder-only" label. KV cache exists to avoid recomputing past K/V during autoregressive generation. In decoder-only models (GPT, Llama) the whole model is the autoregressive decoder, so compression applies throughout. Encoder-decoder models (T5, BART, Whisper) do have a KV cache — in the decoder (both self-attention over generated tokens and cross-attention over fixed encoder outputs) — so KV compression is not impossible there; the encoder simply has no autoregressive cache to compress, and methods like InfoKV / SnapKV are built and tuned for the decoder-only regime (recent-window assumptions, causal structure). So: it's about whether autoregressive decoding is happening, not the architecture name.
Production serving techniques
These belong to different categories — worth keeping straight:
| Technique | Category | Problem it solves |
|---|---|---|
| PagedAttention | memory management (systems) | cache fragmentation / over-reservation |
| Prefix / prompt caching | reuse (systems) | redundant prefill compute / latency |
| KV quantization | precision compression | cache byte footprint |
| Attention sink / sliding window | sparsity / eviction (algorithmic) | cache growing without bound |
| Context-window limit | the constraint itself | max sequence the model supports |
PagedAttention (vLLM) borrows OS virtual-memory paging. Naive contiguous KV allocation reserves max-length per sequence and fragments, capping batch size. PagedAttention splits the cache into fixed-size blocks stored non-contiguously, with a block table mapping logical token positions to physical blocks. Result: near-zero waste (only the last block is partial), cross-sequence sharing (e.g. a shared prompt prefix via copy-on-write), and much higher throughput. It changes where memory lives, not the model math.
Prefix / prompt caching computes a shared prefix's KV once and reuses it across requests (system prompts, few-shot blocks, long documents), cutting prefill compute and time-to-first-token. vLLM's automatic prefix caching and SGLang's RadixAttention (a radix tree of reusable prefixes) are the canonical implementations. This is reuse, not shrinking.
KV quantization stores K/V in low precision (8/4/even 2-bit) instead of fp16/bf16, cutting memory by 2–8×. KIVI (2-bit, per-channel keys, per-token values) and KVQuant are representative. The cost is quantization error accumulating over long sequences — hence reasoning-specific quantizers. Orthogonal to token eviction; stackable.
Attention sink / sliding window. StreamingLLM observed that evicting early tokens for a sliding window collapses quality, because the model dumps surplus attention onto the first few tokens (the "sinks"); the fix is to always keep those few sink tokens plus a recent window, enabling near-infinite streaming with bounded cache. Plain sliding-window attention (Mistral, Longformer) lets each token attend only to the last W positions — bounded cache, O(W) per step, but far-past access is lost except indirectly across layers. (This is the close cousin of R-SWA from Entry 01: window + fixed reference anchor.)
Context-window limit is the maximum sequence length the model supports (positional-encoding range, training length). A hard cap: beyond it you truncate, use retrieval, or extrapolate (RoPE scaling, YaRN). It's the upstream constraint motivating everything above.
Comparison: KV-cache compression methods
| Method | Stage | Scoring signal | Time orientation | Layer budget |
|---|---|---|---|---|
| SnapKV | prefill | attention from a recent observation window (avgpool) | backward | uniform |
| PyramidKV | prefill | same attention scoring as SnapKV | backward | pyramidal (more shallow, less deep) |
| Expected Attention | prefill + decode | expected contribution to future attention (from future-query distribution) | forward (estimated) | uniform |
| InfoKV | prefill + decode | attention + entropy (uncertainty × cross-layer evolution) | backward + information-theoretic | uniform (adaptive optional) |
The lineage: SnapKV (baseline) → PyramidKV (adds the layer axis) → Expected Attention (swaps the backward window for a forward estimate) → InfoKV (adds an information-theoretic signal).
Comparison: attention variants (MHA / MQA / GQA / MLA)
MLA is not GQA. Both shrink the KV cache, by different mechanisms: GQA reduces the number of KV heads (sharing); MLA compresses K/V into a low-rank latent that is up-projected back to per-head K/V at compute time.
| Mechanism | KV heads | Cache / token / layer | How it saves | Quality | Used in |
|---|---|---|---|---|---|
| MHA | = query heads nh | 2·nh·dhead (baseline) | — | baseline | GPT-2/3, Llama-1 |
| MQA | 1 (all queries share) | 2·dhead | full head sharing | noticeable drop | PaLM, Falcon |
| GQA | g groups (1 < g < nh) | 2·g·dhead | grouped head sharing | near-MHA | Llama-2 70B/Llama-3, Mistral, Qwen2 |
| MLA | low-rank latent c (+ small RoPE key) | ≈ dc (+ rope) | low-rank compression + reconstruction | ≈ or > MHA (reported) | DeepSeek-V2/V3 |
Framing: MHA/MQA/GQA/MLA are architectural choices (fixed at training) that change the cache's per-token byte cost; SnapKV / InfoKV and other inference-time token-pruning methods change how many entries (or sequence tokens) are kept. Two orthogonal axes — architecture shrinks per-token cost, compression shrinks count — and they stack.
Explain cross-attention — how does the decoder attend to encoder outputs, and what does "fixed after encoding, doesn't grow" mean? What do fp16 / bf16 stand for? Why is MHA's per-token-per-layer cache 2·nh·dhead? Why does MLA use low-rank, and is it DeepSeek-exclusive? Does "per-token bytes" mean the token's dimension? And is PyramidKV's pyramid (more shallow, less deep) the same as "shallow has more semantic information, deep is more redundant"?
Cross-attention
In an encoder-decoder model, the encoder reads the input once (bidirectional, parallel) and produces a set of output vectors — the encoder outputs / memory (M of them, M = input length), fixed once encoding finishes. Each decoder layer runs masked self-attention → cross-attention → feed-forward. The key fact is where Query, Key, Value come from: in self-attention all three come from the decoder's own tokens; in cross-attention the Query comes from the decoder while the Key and Value come from the encoder outputs.
The decoder holds its query against the encoder's keys (dot product → softmax over input positions → weighted sum of encoder values), which is literally how it reads the input: the decoder is the asker ("what does this output position need from the input?"), the encoder outputs are the memory being read. Because the input doesn't change during generation, cross-attention K/V are computed once per layer at decode start and reused for every generated token — a constant-size cache (= M, the input length). Self-attention's KV cache, by contrast, grows by one entry per generated token. That's what "fixed after encoding, doesn't grow" means.
fp16 / bf16
fp16 = half-precision floating point (IEEE 754 binary16): 1 sign + 5 exponent + 10 mantissa bits. bf16 = bfloat16 / brain floating point (Google Brain): 1 sign + 8 exponent + 7 mantissa bits. bf16 keeps fp32's exponent range (wide dynamic range, fewer overflow/underflow problems) at coarser precision; fp16 has finer precision but a narrow range (needs loss scaling in training). Both are 2 bytes, so they contribute equally to KV-cache size.
Why MHA cache = 2·nh·dhead
Per token, per layer, you cache one Key and one Value. In MHA the Key spans all nh heads at dhead each, so its dimension is nh·dhead; same for Value. K + V = 2·nh·dhead numbers. The "2" is K and V; nh·dhead (the per-token K or V dimension) usually equals dmodel. Total bytes = 2·nh·dhead · L · T · (bytes per number; 2 for fp16/bf16).
Why MLA uses low-rank, and whether it's DeepSeek-exclusive
The KV cache is the memory bottleneck. Rather than caching full per-head K/V (2·nh·dhead), MLA caches a small latent vector c (dimension dc ≪ nh·dhead, a down-projection of the token) and up-projects it back to per-head K and V at attention time. Caching only c saves a lot, yet each head still gets its own reconstructed K/V (unlike GQA's literal sharing), preserving multi-head expressiveness — hence quality ≈ or > MHA. "Low-rank" is the bet that cross-head K/V information compresses into a small latent (Wdown then Wup is a low-rank bottleneck). A RoPE wrinkle — rotary embedding doesn't commute with the absorption trick — is handled by "decoupled RoPE," a small separate key dimension carrying the positional part. As for exclusivity: MLA-the-architecture is DeepSeek's (V2/V3); low-rank-KV-as-a-concept is a broader research line (Palu, MatryoshkaKV, ReCalKV, …). Adoption may have spread since early 2026.
"Per-token bytes" vs "token dimension"
"Per-token bytes" is the memory of one token's KV entry at one layer = (number of K/V values per token) × (bytes per value). "Token dimension" usually means dmodel (the hidden/representation width). They're related but distinct: cache size is driven by the K/V dimension (×2 for K and V), not directly by dmodel. For MHA the KV dimension equals dmodel (so per token = 2·dmodel values); for GQA it's smaller; for MLA it's ≈ dc. Crucially, dmodel (model width) stays the same across MHA/GQA/MLA — what changes is how many K/V numbers you store per token. So "per-token bytes" measures cache footprint, not representation width.
PyramidKV's pyramid vs the redundancy story
PyramidKV gives shallow layers more budget and deep layers less (wide bottom, narrow top). Its stated mechanism is "pyramidal information funneling": attention is broadly spread across many tokens in lower layers and funnels onto a few in higher layers, so lower layers need more retained. This aligns directionally with the redundancy finding (InfoKV §5.2 — deep layers more redundant → less budget), but via a different signal: PyramidKV reads attention dispersion, whereas layer-redundancy analyses read representation convergence (cosine). One wording correction: "shallow has more semantic information" is shaky — the standard interpretability picture is lower layers = more local/surface/syntactic, higher = more abstract; on the redundancy axis, deep layers are more converged/redundant. The precise statement is that shallow representations are more dispersed/diverse (dropping tokens loses more) and deep ones more converged/redundant (safe to drop).