AI Glossary

KV Cache

A key-value (KV) cache stores attention keys and values computed for earlier tokens so an autoregressive transformer can reuse them when generating later tokens. It saves repeated computation but consumes memory.

Also known as: key-value cache

· Updated · Chain of Thought

AI InfrastructureModel Architecture

When a transformer produces another token, it needs information about earlier tokens. Keeping their attention keys and values avoids recomputing that state from scratch. This is runtime state, distinct from the model’s trained weights and from an agent’s persistent memory.

Example: a model’s weights may fit on a GPU, but serving many long conversations can still exhaust its memory because each active request needs cache space. Model architecture, context length, precision, and concurrency all affect that footprint.

Liquid AI’s Maxime Labonne discusses this tradeoff in episode 43, at 3:08: his account compares memory use, inference speed, and quality across architectures. That is a concrete reason to benchmark your serving workload rather than compare parameter counts alone. Cache quantization or offloading can change memory use, with their own quality or performance tradeoffs.

Reuse past attention keys and valuesAfter the prompt A B is processed, its keys and values are cached and token C is predicted. A decode step consumes C, computes its query, key and value, and attends to the cached A B state plus C. The result predicts D, while the cache now holds A B C. The weights do not change. Reuse past attention keys and valuesCausal autoregressive generation: consume the latest token to predict the next Prompt A B → predict CCache holds keys and values for A and BConsume token CCompute Qc, Kc, VcAttention for CUse K/V from A, B, CPredict DNext tokenCache after this step: K/V for A, B and C; model weights unchangedReuse avoids recomputing past K/V; each active request still needs cache memory.
The token sequence is illustrative. Hugging Face’s cache explanation describes reusing past keys and values while adding the current token’s state. Download the image

Sources

Go deeper

From the conversation