A local model can load its weights and still run out of memory when you raise the context limit. The KV cache is often what grew, even though the model size stayed the same.
The short answer: A KV cache holds the attention keys and values already computed for tokens in the active sequence, so the model can reuse them while generating the next token. Its size is approximately 2 × layers × KV heads × head dimension × context tokens × bytes per value; multiply again for concurrent sequences. Lower num_ctx first if you do not need the full window, then try cache quantization. Prefix caching is a separate optimization that reuses a matching prompt prefix across requests.
What is a KV cache, and what does it store?
A KV cache stores the key and value vectors produced by each attention layer for tokens already processed. On the next generation step, the model adds the new token's K/V values and uses the prior entries instead of projecting those past tokens again; it doesn't cache the query vector. Hugging Face’s cache explanation describes the per-layer K/V reuse.
That makes the cache grow with sequence length — a longer prompt and a longer generated answer both occupy the context window, so the same model can need different cache memory at 8K and 32K. The cache is separate from model weights: quantizing weights doesn't automatically quantize the KV cache, as Hugging Face’s cache-quantization guide discusses them as separate techniques.
How do I calculate KV cache memory?
Use the model's KV-head count, not its total attention-head count, and count both K and V. For one sequence, the estimate is:
cache bytes = 2 × layers × num_key_value_heads × head_dim × context tokens × bytes per valueThe leading 2 is for keys and values. With grouped-query attention, num_key_value_heads can be much smaller than num_attention_heads; use the field in the model config. Hugging Face’s KV-cache quantization guide gives the formula with num_key_value_heads; its memory example shows the equivalent layer × head × dimension calculation, and its cache guide describes the per-layer tensors.
For a quick check, Qwen3-8B has 36 layers, 8 KV heads, and a 128-wide head in its Hugging Face config. At FP16, that's 2 × 36 × 8 × 128 × 2 = 147,456 bytes per token, or about 1.12 GiB at 8,192 tokens. Two simultaneous 8K sequences need about twice that cache storage before runtime overhead.
The figures below use GiB (2³⁰ bytes) and one sequence. For q8_0 and q4_0, they include each 32-value block's scale: ggml’s q8_0/q4_0 definitions make those 34 and 18 bytes per 32 values, rather than exactly 1 and 0.5 bytes per value. The calculation script and its inputs are listed in the evidence log.
How much KV cache do common local models need?
These are calculated cache sizes from each model's published config, not measured VRAM. Each context cell lists FP16 / q8_0 / q4_0 GiB for one sequence.
Dense models (every layer is full attention):
| Model and config fields | Cache per token (FP16) | 8K context | 32K context | 128K context |
|---|---|---|---|---|
| Llama 3.1 8B · 32 layers, 8 KV heads, head dim 128 | 128 KiB | 1.00 / 0.53 / 0.28 | 4.00 / 2.12 / 1.12 | 16.00 / 8.50 / 4.50 |
| Qwen3-8B · 36 layers, 8 KV heads, head dim 128 | 144 KiB | 1.12 / 0.60 / 0.32 | 4.50 / 2.39 / 1.27 | 18.00 / 9.56 / 5.06 |
| Mistral Small 3.2 24B · 40 layers, 8 KV heads, head dim 128 | 160 KiB | 1.25 / 0.66 / 0.35 | 5.00 / 2.66 / 1.41 | 20.00 / 10.62 / 5.62 |
| Qwen3-30B-A3B · 48 layers, 4 KV heads, head dim 128 | 96 KiB | 0.75 / 0.40 / 0.21 | 3.00 / 1.59 / 0.84 | 12.00 / 6.38 / 3.38 |
| Qwen3-32B · 64 layers, 8 KV heads, head dim 128 | 256 KiB | 2.00 / 1.06 / 0.56 | 8.00 / 4.25 / 2.25 | 32.00 / 17.00 / 9.00 |
| Llama 3.3 70B · 80 layers, 8 KV heads, head dim 128 | 320 KiB | 2.50 / 1.33 / 0.70 | 10.00 / 5.31 / 2.81 | 40.00 / 21.25 / 11.25 |
Hybrid and sliding-window models (only some layers grow with context):
| Model and config fields | Growing cache per token (FP16) | 8K context | 32K context | 128K context |
|---|---|---|---|---|
| gpt-oss-20b · 12 full-attention + 12 sliding-window (128-token) layers, 8 KV heads, head dim 64 | 24 KiB | 0.19 / 0.10 / 0.05 | 0.75 / 0.40 / 0.21 | 3.00 / 1.60 / 0.84 |
| Qwen3.5-9B · 8 full-attention + 24 Gated DeltaNet layers, 4 KV heads, head dim 256 | 32 KiB | 0.25 / 0.13 / 0.07 | 1.00 / 0.53 / 0.28 | 4.00 / 2.12 / 1.12 |
| Qwen3.5-27B · 16 full-attention + 48 Gated DeltaNet layers, 4 KV heads, head dim 256 | 64 KiB | 0.50 / 0.27 / 0.14 | 2.00 / 1.06 / 0.56 | 8.00 / 4.25 / 2.25 |
The parameter count is a weak predictor. Qwen3-30B-A3B is a mixture-of-experts model with 30B total parameters, yet it needs less cache per token than Qwen3-8B because it has 4 KV heads instead of 8.
Architecture matters more than size. At 128K, Qwen3.5-9B's growing cache is 4 GiB against 16 GiB for Llama 3.1 8B, because only 8 of its 32 layers are full attention: the Qwen3.5-9B model card gives the layout as "8 × (3 × (Gated DeltaNet → FFN) → 1 × (Gated Attention → FFN))". The Gated DeltaNet paper (ICLR 2025) describes that layer family as a linear RNN "with matrix-valued states", a fixed-size state rather than a per-token cache, so the table leaves it out.
gpt-oss-20b's sliding-window layers stop at 128 tokens; Hugging Face's cache docs say such a cache "will stop growing" once those layers reach the window size. Whether a runner allocates this little depends on its support for the architecture, so compare with the memory it reports after loading.
The 128K column is arithmetic, not a promise that the model or runner accepts that length. The Qwen3 configs list max_position_embeddings: 40960; the Llama, Mistral Small and gpt-oss configs list 131,072; the Qwen3.5 configs list 262,144. Check the runner's model metadata before raising the context.
To turn a table cell into a rough device budget, add the model's weight memory, runtime buffers, activations, and any other loaded models. Then multiply the one-sequence cache by the number of active sequences. Ollama documents that parallel requests reserve context memory per request, so a 32K setting with two parallel sequences can require about twice the per-sequence cache.
Which settings reduce or reuse the cache?
Lower the context length when you don't need the full model window; use KV cache quantization when you need that length but lack memory. Prefix caching helps with repeated prefixes across requests, but it doesn't make the active sequence's cache formula smaller. The settings below were checked in official documentation on October 4, 2026; verify support in the runtime version you have installed.
| Runtime | Setting to change | What it does and quality cost |
|---|---|---|
| Ollama | num_ctx; OLLAMA_FLASH_ATTENTION=1; OLLAMA_KV_CACHE_TYPE=q8_0 or q4_0 | Set num_ctx per run/request to cap the context. Flash Attention is automatically used when the backend supports it; the environment variable forces it on. Quantized K/V requires Flash Attention and is a global setting. Ollama says q8_0 uses about half the FP16 memory with usually no noticeable quality impact; q4_0 uses about one quarter, with small-to-medium loss that can be more noticeable at long contexts. See Ollama’s FAQ. |
| llama.cpp | -c 32768 -fa on --cache-type-k q8_0 --cache-type-v q8_0 | -c/--ctx-size sets context; -fa/--flash-attn enables Flash Attention; --cache-type-k and --cache-type-v set each cache’s type. The documented defaults are model-loaded context (-c 0), Flash Attention auto, and FP16 K/V. The docs give no universal quality-loss figure for q8_0/q4_0; evaluate your model and task. See the llama.cpp server options. |
| LM Studio | K/V cache quantization for a llama.cpp model; in code, llamaKCacheQuantizationType and llamaVCacheQuantizationType in the load config | The LM Studio 0.3.7 notes added K/V cache quantization for llama.cpp models (llama.cpp runtime 1.9.0+); the load-config docs name the key and value properties. Its docs say lower precision saves memory and may affect output quality; V-cache quantization requires Flash Attention. |
| vLLM | --max-model-len 32768 --kv-cache-dtype fp8 --enable-prefix-caching --gpu-memory-utilization 0.90 | --max-model-len caps prompt plus output; --kv-cache-dtype selects cache precision; prefix caching is enabled by default in the docs checked; --gpu-memory-utilization controls the instance’s device-memory budget (default 0.92), not bytes per token. FP8 support depends on backend/hardware. The vLLM CLI reference does not give a general quality-loss number for FP8; benchmark your model/task before keeping it. |
For Ollama, the docs show /set parameter num_ctx 4096 in an interactive run and options.num_ctx in the API. Replace the sample length with a context your model supports and your memory can hold. For llama.cpp, this is a server command shape; model.gguf is the local file you supply:
llama-server -m model.gguf -c 32768 -fa on --cache-type-k q8_0 --cache-type-v q8_0For the vLLM KV cache settings, the flag form above is copyable with an actual model ID, for example:
vllm serve Qwen/Qwen3-8B --max-model-len 32768 --kv-cache-dtype fp8 --enable-prefix-caching --gpu-memory-utilization 0.90These commands are docs-only; I didn't run a model server in this research session. Prefer the context reduction first, before any KV cache compression, because it keeps the original cache precision. If the measured task quality changes after quantizing the cache, return to FP16 or try q8_0 before q4_0. Running a coding agent locally with Ollama shows the local-provider setup around Ollama.
Why is cached input cheaper with an API provider?
Provider prompt caching can reuse a matching prompt prefix from an earlier request, so the provider does less repeated prefill work and charges a lower cached-input rate. Both big providers say K/V state is what they keep: OpenAI's docs say its prompt cache "stores key-value (KV) tensors, not the tokens themselves", and Anthropic's say "KV (key-value) cache representations" of cached content are held in memory only. Read each provider's current pricing and cache rules before estimating a bill; our prompt-caching explainer covers the billing distinction.
For a better chance of reuse, keep stable system instructions and tool definitions at the start of the prompt, then append changing user content. If the shared prefix changes, later tokens cannot reuse the old matching prefix under OpenAI's documented rules.
Local KV cache quantization changes how much memory one active sequence uses. Provider prompt caching reuses a prefix between requests and only helps when that prefix matches under the provider's rules. See OpenAI’s prompt-caching guide and Anthropic’s prompt-caching docs.
What does this estimate leave out?
The table estimates tensor storage for one request's K/V entries; it isn't a GPU allocation measurement or a promise that a given card will run the model. It excludes weights, activations, temporary attention workspace, allocator/runtime overhead, multimodal payloads, other requests, and the fixed recurrent state of Gated DeltaNet layers. MLA models such as DeepSeek's compress the K/V cache into a latent vector (DeepSeek-V2 paper, 2024), a different geometry, so I left them out rather than applying this equation to them.
A useful check is to compare the calculated per-sequence cache with the runtime's actual memory after load, then multiply for intended concurrency and leave room for everything the formula omits. Octomind is ours — its local: provider can point to a local Ollama server, and its cache-aware compaction accounts separately for cache reads and writes; those behaviors are documented in our local Ollama guide and prompt-caching explainer.
Get Octomind — run an agent against a local Ollama model.
FAQ
Does a KV cache store the prompt text?
No. It stores the key and value tensors computed from processed tokens, not a copy of the prompt as plain text.
Does a longer context always use more cache memory?
For ordinary full attention, yes: the cache grows with the number of tokens retained, along with layers, KV heads, head dimension, and bytes per value. Sliding-window, hybrid, and MLA architectures need a different calculation.
Is q8_0 the same as FP8?
No. q8_0 is a llama.cpp/GGML block quantization format, while FP8 is a floating-point cache type supported by particular serving backends. Similar bit width doesn't make their layouts, quality behavior, or compatibility interchangeable.
Does prefix caching lower the memory for one long conversation?
Prefix caching reuses matching prompt work across requests; it isn't a substitute for setting the active context length or reducing that sequence's K/V precision. In a serving engine it may improve repeated-prefill cost, but cache blocks still consume the engine's cache capacity.



