The KV cache and three ways to shrink it
Barakaeli Lawuo, Jun 24, 2026
In May 2024 DeepSeek-AI reported that its new model, DeepSeek-V2, needs 93.3% less KV cache than DeepSeek 67B, the dense model it replaced. Over the same comparison it cut training cost by 42.5% and raised maximum generation throughput to 5.76 times the old figure.1 On a single node of 8 H800 GPUs, the paper says, DeepSeek-V2 generates more than 50,000 tokens per second.1
The throughput gain follows from the memory number. A server that stores less per token can keep more requests in flight at once, and the paper says as much: the smaller cache is what lets the deployed model "serve a much larger batch size."1 This post follows that budget as it shrinks. First comes why the cache exists at all, then how big it is, then three changes to attention that each made it smaller: multi-query attention in 2019, grouped-query attention in 2023, and DeepSeek's multi-head latent attention in 2024. Each one bought memory and paid for it somewhere else.
Generating without a cache redoes the whole prefix
A language model writes one token (a word or piece of a word) at a time. Each new token is fed back in as input for the next step, which the Transformer paper calls being auto-regressive.2 Training does not have this problem, because the full target text is known in advance and every position can be processed in parallel. At generation time, the output at one position decides the input at the next, so the steps must run in order.3
Inside every layer, each token is turned into a query, a key and a value (defined properly in the next section). The key and value for a past token never change once that token is written, because later tokens cannot influence earlier ones in a decoder. A naive generator ignores this and runs the whole prefix through the model again at every step.
The Transformer paper lists the cost of one self-attention layer over a sequence of length with width as .2 Rerunning the prefix means paying that price again at every step , plus for the projections. Summing over the steps (standard arithmetic, not a figure from the papers) gives roughly for the whole output. If instead the model stores each token's key and value the first time it computes them, step only has to project one new token and compare one query against stored keys, which is per layer. That totals . Those stored tensors are the KV cache.
Shazeer's 2019 paper analyses exactly this incremental loop. Under his simplifying assumptions, the arithmetic across steps is for a batch of sequences, the same as processing them in parallel. The memory traffic is the problem: , where the first term is loading the cached keys and values and the second is loading the weights.3 The ratio of memory access to arithmetic is . When that ratio is near 1, he writes, memory bandwidth becomes a major bottleneck on modern hardware, where compute can be two orders of magnitude faster than memory.3 A bigger batch shrinks , memory permitting. The term comes from reloading the cache at every step, and batching does nothing about it.
So the cache removes the repeated arithmetic but leaves a new cost behind. It has to be stored, and it has to be read in full once per generated token.
The attention equation, one symbol at a time
holds the queries, one row per position that is asking a question. holds the keys, one row per position that can be looked at, and holds the values, the content that gets passed along. Each row of is a list of dot products between one query and every key, a score for how well each earlier token matches. Dividing by , where is the key width, keeps those scores from growing large as vectors get wider; the authors suspect large scores push the softmax into regions with tiny gradients.2 The softmax turns each row of scores into weights that sum to 1, and multiplying by returns a weighted average of the values.2 During generation, a mask sets the scores for future positions to minus infinity so a token can only look backward.2
Multi-head attention runs of these in parallel. Each head has its own learned projections that produce its own queries, keys and values from the same input, and the head outputs are concatenated and projected again.2 The base Transformer used heads of width .2 For the cache, the fact that matters is that every head carries its own keys and values, so all of them must be stored.
- KV cache
- the keys and values already computed for every earlier token, in every layer and head, kept in accelerator memory so they are not recomputed at each step.
- Decode step
- one pass that produces one new token. It reads the weights and the entire KV cache once.
- Memory bandwidth
- how fast data can move from the accelerator's main memory to its compute units. Decoding is usually limited by this, not by arithmetic.
- Head
- one of the parallel attention computations inside a layer, each with its own projections.
Sizing the cache
DeepSeek-V2's paper states the count for standard multi-head attention: it caches elements per token, where is the number of heads, the width of each head and the number of layers.1 The factor 2 is one key and one value. Multiplying out to bytes for a whole batch is standard arithmetic:
Term by term: layers each keep their own cache. is how many distinct key and value heads a layer stores; in plain multi-head attention it equals the number of query heads. is the width of each head. is bytes per element, 2 for 16-bit numbers. is the number of tokens in the context, and is the number of sequences served together. Every term is a straight multiplier, and the last two grow with traffic.
A worked example uses DeepSeek 67B, the model in the opening. Its paper lists 95 layers, a model width of 8,192, 64 query heads, 8 key/value heads and a 4,096-token context.6 That gives a head width of 128 and elements per token. At 2 bytes each, that is about 389 KB per token and roughly 1.6 GB for one full 4,096-token sequence. (This arithmetic is mine, not a number printed in either paper.)
At larger scale the figures get big. Pope and colleagues at Google give one: for a 500B+ parameter model with multi-head attention, a batch of 512 and a context of 2,048, the KV cache totals 3 TB, three times the size of the model's weights. The chip must load that cache from off-chip memory once for every token generated, and the compute cores sit mostly idle while it does.4 They also report that at batch sizes and sequence lengths around 512 and 2,048 and above, the time to load the cache dominates the time to load the weights.4
Only a few terms in the size formula are open to change. Context length and batch size are what users want, so they are left alone. Precision can be lowered, and that is the subject of quantization work. The three designs below attack .
Multi-query attention: every query head shares one key and one value
Noam Shazeer's 2019 proposal is a small edit. Multi-query attention (MQA) keeps many query heads but has all of them share a single set of keys and values.3 In the size formula, drops from to 1. His analysis shows the troublesome memory-to-arithmetic term falls from to , a factor of .3
He tested it on WMT 2014 English to German translation with a 6-layer encoder-decoder model with , 8 heads of width 128 and 211 million parameters. To keep the parameter count equal, the MQA model's feed-forward layers were widened from 4,096 to 5,440.3 Incremental greedy decoding of 1,024 sequences on one TPUv2 took 46 microseconds per output token in the decoder for the baseline and 3.8 for MQA, about 12 times faster. With beam search of width 4 the decoder cost fell from 203 to 32 microseconds per token.3
Quality moved a little. Dev-set BLEU went from 26.7 to 26.5, and log perplexity rose from 1.424 to 1.439. On the test set with beam search, MQA actually scored the higher BLEU, 28.5 against 28.4.3 On the Billion-Word language modeling benchmark, dev perplexity went from 29.9 to 30.2.3 Shazeer also tried the obvious alternative of shrinking the cache by using fewer or narrower heads. Those models were worse than MQA on both tasks; one head of width 128 dropped translation BLEU to 25.8.3 Sharing keys and values cost less quality than removing them.
Grouped-query attention: a few key heads instead of one
Four years later, Joshua Ainslie and colleagues at Google Research gave two reasons MQA had not spread everywhere. It can lower quality and make training unstable, and it may not be feasible to train a separate model just for fast inference. Many public models, including T5 and LLaMA, had been trained with full multi-head attention.5
Their answer has two parts. Grouped-query attention (GQA) splits the query heads into groups, and each group shares one key head and one value head. GQA-1 is MQA and GQA- is ordinary multi-head attention.5 The second part is uptraining: take an existing multi-head checkpoint, build each group's key and value projections by averaging (mean pooling) the original heads in that group, and continue pre-training for a small fraction of the original steps.5 Mean pooling beat both keeping the first head and starting the shared heads from random weights.5 For the extra training took about 600 TPUv3 chip-days.5

The authors also argue that grouping scales better than MQA. Larger models tend to have more heads, so cutting to one key head is a harsher cut the bigger the model gets, while GQA keeps the reduction proportional. And standard sharding across chips copies MQA's single key/value head onto every partition, which GQA avoids.5
They measured on T5 XXL, uptrained with , across summarization, translation and question answering. GQA was applied to decoder attention only, since the encoder runs in parallel and is not bandwidth bound.5
Read together, the two charts show the trade. Uptrained MQA-XXL ran about 6.3 times faster than multi-head XXL and lost 0.6 points of average score. GQA with 8 groups ran about 5.4 times faster and lost 0.1.5 Both XXL variants were faster than the much smaller MHA-Large and scored higher.5 Going from 1 group to 8 added only modest time, with the cost climbing as groups approached full multi-head, so the authors chose 8 as a middle ground.5 GQA also worked reasonably straight after conversion, while MQA needed the uptraining to be useful.5
A later ablation in the DeepSeek-V2 paper tells a less comfortable story. There, 7B dense models trained from scratch on 1.33 trillion tokens were compared on harder benchmarks. MMLU accuracy was 45.2 with multi-head attention, 41.2 with 8-group GQA and 37.9 with MQA.1 My reading is that the settings differ enough (decoder-only models trained from scratch instead of uptrained encoder-decoders, and multiple-choice knowledge tests instead of summarization) that the two papers do not contradict each other. They do show that "close to multi-head" depends on what you measure.
Multi-head latent attention: cache a compressed vector instead
DeepSeek's position is that MQA and GQA cache less but "their performance does not match MHA."1 Its multi-head latent attention (MLA) keeps full multi-head keys and values during computation and changes what gets stored. Each token's hidden state is projected down to one short latent vector, and the keys and values for every head are rebuilt from it:
is the down-projection, a learned matrix that maps the hidden state to the latent of width . and are up-projections that expand it back to full-width keys and values. Only is cached, so a layer stores elements per token instead of .1 The up-projection for keys can be folded into the query projection, and the one for values into the output projection, so at inference the full keys and values never have to be built at all.1 Queries get a similar compression, but only to save activation memory in training; it does not touch the cache.1
One complication shows the constraint the design works under. DeepSeek uses rotary position embeddings (RoPE), which rotate queries and keys by an amount that depends on position. Applied to the compressed keys, that rotation would sit between the query projection and , the fold would no longer work, and the paper says the model would have to "recompute the keys for all the prefix tokens during inference", which is the waste the cache exists to avoid.1 The fix is a separate small key of width , shared across heads, that carries the position information and is cached alongside the latent. The total becomes elements per token.1
DeepSeek-V2 uses 60 layers, 128 heads of width 128, and .1 Per layer that is 576 cached elements where full multi-head attention with those heads would store 32,768. The paper puts it as equal to GQA with only 2.25 groups.1 Its measured comparison is in Table 9, MoE models that differ only in attention:
Quality went the other way from MQA. For the large models, MLA scored 59.0 on MMLU against 57.5 for multi-head attention and 50.7 against 46.6 on BBH. For the small models MLA was ahead on three of four benchmarks and behind on C-Eval, 50.9 against 51.6.1 The large pair was trained on only 420 billion tokens.1
Now the opening number can be traced. DeepSeek-V2 caches elements per token, which matches the 34.6K in Table 9. The paper also says the deployed model was converted to FP8 and its KV cache quantized to about 6 bits per element on average.1 My arithmetic, assuming DeepSeek 67B stored its 194,560 elements per token in 16 bits: at 16 bits DeepSeek-V2's cache would be about 82% smaller, and at 6 bits it is 25,920 bytes against 389,120, a reduction of 93.3%. The paper does not show this breakdown, so treat it as a reconstruction. It suggests the headline combines two savings, a change to the architecture and a change to precision.
What the papers say they could not show
Each paper is candid about its gaps. Shazeer's timing runs used fixed-shape tensors, so the cache was padded to the full 128 positions and every decoding step took the same time; a cache that grew with the sequence could have been faster early on.3 DeepSeek's decoupled RoPE key is a second cached tensor, needed only because the compression trick and rotary positions do not combine.1
The GQA authors state the widest limit. The memory overhead they target matters most for long generations, and long outputs are hard to evaluate. They scored summarization with ROUGE, which they call "a flawed evaluation that does not tell the whole story," and conclude that "it is difficult to be certain our trade-offs are correct."5 They did not compare their uptrained model to one trained from scratch, and they tested only encoder-decoder models, though they expect GQA's advantage over MQA to be larger in decoder-only models.5 In the same paragraph they note that decoder-only models had become extremely popular.5 That is the setting where GQA is now most used, and its own paper never measured it.
Sources
- DeepSeek-AI (2024). DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model. arXiv:2405.04434
- Vaswani et al. (2017). Attention Is All You Need. arXiv:1706.03762
- Shazeer (2019). Fast Transformer Decoding: One Write-Head is All You Need. arXiv:1911.02150
- Pope et al. (2022). Efficiently Scaling Transformer Inference. arXiv:2211.05102
- Ainslie et al. (2023). GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints. arXiv:2305.13245
- DeepSeek-AI (2024). DeepSeek LLM: Scaling Open-Source Language Models with Longtermism. arXiv:2401.02954