Back to Blog
September 30, 2026

The short version: The KV cache for one 128K-token sequence needs 16 GiB for Llama 3.1 8B and 40 GiB for Llama 3.1 70B, at 16-bit precision. The formula is 2 × layers × KV heads × head dimension × bytes per value × tokens. Grouped-query attention, FP8 or 2-bit cache quantization, and latent attention can cut that by 2x to 16x. Serving several long requests at once multiplies the total.

What is the formula for KV cache memory?

During generation, a transformer stores a key vector and a value vector for every token, in every layer, for every key-value head. That store is the KV cache. It lets the model attend to earlier tokens without recomputing them.

The size per token is:

bytes per token = 2 × n_layers × n_kv_heads × head_dim × bytes_per_value

The 2 counts keys and values. Multiply by the number of tokens in the sequence, then by the number of sequences in the batch. At BF16 or FP16, bytes_per_value is 2.

The vLLM PagedAttention paper gives a worked example. For the 13B-parameter OPT model, one token needs 800 KB: 2 (keys and values) × 5120 (hidden size) × 40 (layers) × 2 (bytes per FP16). OPT uses full multi-head attention, so the hidden size equals KV heads times head dimension. The paper notes that OPT's 2,048-token limit capped a single request at about 1.6 GB. At 128K tokens, the same arithmetic gives numbers that exceed most single GPUs.

How much VRAM does a 128K context need for common models?

Here "128K" means 131,072 tokens, the max_position_embeddings value in the Llama 3.1 and Mistral NeMo configs. All figures assume a BF16 cache and a batch of one. The architecture numbers come from each model's official config.

  • Llama 3.1 8B (32 layers, 8 KV heads, head dim 128, per Meta's llama-models repository): 128 KiB per token, 16 GiB at 128K.
  • Llama 3.1 70B (80 layers, 8 KV heads, head dim 128): 320 KiB per token, 40 GiB at 128K.
  • Llama 3.1 405B (126 layers, 8 KV heads, head dim 128): 504 KiB per token, 63 GiB at 128K.
  • Mistral NeMo 12B (40 layers, 8 KV heads, head dim 128, per its Hugging Face config): 160 KiB per token, 20 GiB at 128K.
  • Phi-3-mini-128k (32 layers, 32 KV heads, head dim 96, per its Hugging Face config): 384 KiB per token, 48 GiB at 128K.

The Phi-3 line is the surprise. It is a 3.8B-parameter model whose BF16 weights take roughly 7.6 GB. Because it uses full multi-head attention, its 128K cache is about six times larger than its weights. Llama 3.1 8B is a bigger model with a cache one third the size.

These numbers cover the cache only. Add the model weights, activations, and framework overhead. Serving engines also reserve memory in advance, so real usage runs higher than the raw formula.

Why do KV heads matter more than parameter count?

The cache scales with KV heads, not query heads. The GQA paper (Ainslie et al., 2023) introduced grouped-query attention, where several query heads share one key-value head. It sits between full multi-head attention and multi-query attention, which uses a single KV head. The authors report quality close to multi-head attention at speed comparable to multi-query attention.

Llama 3.1 8B has 32 query heads but only 8 KV heads. If it used 32 KV heads, its 128K cache would be 64 GiB instead of 16 GiB. That 4x saving is why nearly every recent open model ships with GQA.

DeepSeek-V2 goes further with Multi-head Latent Attention (MLA), which compresses keys and values into a latent vector. The paper reports a 93.3% smaller KV cache than DeepSeek 67B, and a 128K context length.

How much does quantizing the KV cache save?

Bytes per value is the easiest term to change. An FP8 cache halves every figure above, so Llama 3.1 8B at 128K drops from 16 GiB to 8 GiB. Most serving engines expose this as a setting.

Research pushes lower. KIVI (Liu et al., 2024) quantizes the cache to 2 bits: keys per channel, values per token. The authors report that Llama, Falcon, and Mistral models keep almost the same quality with 2.6x less peak memory, including weights. That enabled up to 4x larger batches and 2.35x to 3.47x higher throughput. At a pure 2 bits per value, the Llama 3.1 8B cache at 128K would be about 2 GiB before quantization metadata.

Why does batch size change the answer?

The formula is per sequence. Eight concurrent 128K requests to Llama 3.1 70B need 320 GiB of BF16 cache, which is more than four 80 GB GPUs hold before any weights load. The PagedAttention paper frames serving throughput as memory-bound for this reason. vLLM's paged allocation reduces waste from fragmentation, and the authors report 2x to 4x throughput gains over earlier systems. It cannot shrink the cache a request actually uses.

This is also why long prompts are expensive even when the model accepts them. Every token you keep in context holds memory for the life of the request. Prompt caching can reuse a shared prefix across requests, but the prefix still occupies cache while in use.

How do RLMs change the memory math?

A Recursive Language Model (Zhang, Kraska, and Khattab) treats a long prompt as part of an external environment. The model examines and decomposes the prompt in code, then calls itself recursively over snippets. The paper reports that RLMs process inputs up to two orders of magnitude beyond the model's context window.

In memory terms, the full document lives in the REPL as ordinary data, not in the KV cache. Each root or sub-call holds only its own working context. A 10-million-token corpus never needs a 10-million-token cache. Peak VRAM depends on the largest single call, plus how many sub-calls run at once. For that reason RLMs pair well with small models on long inputs. The tradeoff is more calls and more latency, which we cover separately.

The bottom line

Use 2 × layers × KV heads × head dim × bytes × tokens, with the numbers from the model's config. At BF16 and 131,072 tokens, expect about 16 GiB for Llama 3.1 8B, 20 GiB for Mistral NeMo 12B, and 40 GiB for Llama 3.1 70B, per sequence. Check the KV head count first: a small model without GQA can need more cache than a much larger one with it. FP8 halves the bill, and 2-bit methods such as KIVI go further.

References & Further Reading

  1. Kwon, W., et al. "Efficient Memory Management for Large Language Model Serving with PagedAttention." SOSP 2023, arXiv:2309.06180. Source of the 800 KB per token OPT-13B worked example and the memory-bound serving argument. arxiv.org/abs/2309.06180
  2. Ainslie, J., et al. "GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints." arXiv:2305.13245, May 2023. Defines grouped-query attention. arxiv.org/abs/2305.13245
  3. DeepSeek-AI. "DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model." arXiv:2405.04434, May 2024. Reports a 93.3% smaller KV cache with Multi-head Latent Attention. arxiv.org/abs/2405.04434
  4. Liu, Z., et al. "KIVI: A Tuning-Free Asymmetric 2bit Quantization for KV Cache." arXiv:2402.02750, February 2024. Reports 2.6x less peak memory and 2.35x to 3.47x throughput. arxiv.org/abs/2402.02750
  5. Zhang, A. L., Kraska, T., Khattab, O. "Recursive Language Models." arXiv:2512.24601, December 2025. Treats long prompts as an external environment and processes inputs far beyond the context window. arxiv.org/abs/2512.24601
  6. Meta. llama-models repository, models/sku_list.py. Architecture arguments for Llama 3.1 8B, 70B and 405B (layers, heads, KV heads). github.com/meta-llama/llama-models/blob/main/models/sku_list.py
  7. Mistral AI. Mistral-Nemo-Instruct-2407 config.json on Hugging Face. Layers, KV heads, head dimension and 131,072 max positions. huggingface.co/mistralai/Mistral-Nemo-Instruct-2407/blob/main/config.json
  8. Microsoft. Phi-3-mini-128k-instruct config.json on Hugging Face. 32 layers with 32 KV heads (no GQA) and 131,072 max positions. huggingface.co/microsoft/Phi-3-mini-128k-instruct/blob/main/config.json
FAQ

Frequently asked questions

Is 128K context 128,000 or 131,072 tokens?

Model configs such as Llama 3.1 and Mistral NeMo set max_position_embeddings to 131,072, which is 128 x 1,024. At exactly 128,000 tokens the cache is about 2% smaller. For Llama 3.1 8B that is about 15.6 GiB instead of 16 GiB.

Does the KV cache include the prompt or only generated tokens?

Both. Every token in the sequence, prompt and output, adds keys and values in every layer. A 120K-token prompt with an 8K-token answer fills the same cache as a 128K sequence.

Why is my real VRAM use higher than the formula?

The formula covers the cache only. Model weights, activations, CUDA context and the serving engine's pre-allocated memory all add to it. Engines such as vLLM reserve a fraction of GPU memory up front for the cache.

Does FP8 KV cache hurt quality?

It depends on the model and the engine's scaling method, so test it on your own tasks. Research on 2-bit caches such as KIVI reports almost the same quality on Llama, Falcon and Mistral models, which suggests 8-bit usually has headroom. Measure before relying on it for accuracy-critical work.

Do RLMs avoid the KV cache entirely?

No. Every root call and sub-call still uses a KV cache for its own context. The difference is that the full input sits in the REPL as data, so no single call has to hold the whole document in cache.