본문으로 건너뛰기
L14.6

KV Cache and Generation Throughput

Goal

Explain why autoregressive serving reuses KV-cache state, how cache memory scales with active tokens, and why scheduler admission affects throughput.

Reuse past attention work​

Imagine reading a long story and trying to predict the next word. If every prediction forced you to close the book, reopen page 1, and reread everything you had already seen, most of your work would be repeated. It is faster to keep useful notes from the earlier pages and reuse them.

An autoregressive language model has a similar repeated-work problem because it produces one next token at a time. Remember from L5.3 — Queries, Keys, and Values: attention creates key and value information for tokens. During serving, earlier tokens have already produced that information. The KV cache keeps that information from previous positions so later decoding steps can reuse it instead of rebuilding all of it from the beginning.

The cache is therefore a speed-for-memory tradeoff. Reuse saves compute, but the notes take space. A request with 100 cached tokens needs less cache state than one with 8,000 cached tokens, and several active requests keep their own state at the same time.

Cache memory becomes a service constraint​

That speed-for-memory tradeoff becomes a service constraint: longer prompts, longer outputs, and more concurrent requests all keep more cached token state alive at the same time.

request A: 100 cached tokens → small cache footprint
request B: 8000 cached tokens → much larger cache footprint
many active requests keep separate cache state

KV cache growth: more active tokens keep more attention state in memory

Move the cached-token count while keeping the toy architecture fixed. The memory bar grows with the number of active cached tokens, making the speed-for-memory tradeoff visible.

100 cached tokens · about 12.5 MiB
Cached KV state1008kNow: 100 ≈ 12.5 MiB

KV bytes ≈ cached tokens × layers × 2 (Key + Value) × KV heads × head dimension × bytes per element

Under these assumptions, each cached token adds about 128 KiB of KV state, so 8,000 tokens use about 80× as much cache memory as 100 tokens.

AssumptionValue
Layers32
KV heads8
Head dimension128
Bytes per element2
Cached tokensApproximate KV memoryRelative to 100 tokens
10012.5 MiB1×
1,000125.0 MiB10×
8,0001000.0 MiB80×

The exact cache size depends on model architecture, precision, parallelism, and runtime layout. You do not need one universal bytes-per-token number. Production capacity planning should measure or obtain the runtime's cache accounting for the exact deployed model.

Go deeper: estimate KV-cache memory

For a decoder-only transformer, one useful approximation is:

KV bytes
≈ cached_tokens
× layers
× 2 # key + value
× kv_heads
× head_dim
× bytes_per_element

Assume a simplified model with:

layers = 32
kv_heads = 8
head_dim = 128
bytes_per_element = 2 # for example, BF16

Then the approximate cache per token is:

32 × 2 × 8 × 128 × 2
= 131,072 bytes
= 128 KiB per cached token

That connects directly to the earlier intuition:

100 cached tokens ≈ 12.5 MiB
8,000 cached tokens ≈ 1,000 MiB ≈ 0.98 GiB

Eight simultaneous 8,000-token requests would therefore need about 7.8 GiB of KV state under these assumptions, before allocator overhead and other runtime memory.

This is an estimate, not a universal formula for every deployment. Grouped-query or multi-query attention changes the number of KV heads. KV quantization changes bytes per element. Tensor parallelism can shard cache state. Paged allocators, block rounding, prefix sharing, and runtime metadata also change actual device memory. Use this calculation to build intuition, then confirm the deployed runtime's accounting.

Let the runtime manage cache blocks​

Serving runtimes actively manage this memory. vLLM documents PagedAttention and explicit KV-cache controls among its serving features. Its current documentation also exposes configuration for GPU memory utilization and KV-cache memory. The specific defaults can change, so treat them as deployment configuration rather than course constants.

Throughput and cache capacity interact. Admitting more active sequences can keep the accelerator busy, but high cache pressure can cause eviction (removing cached state to free memory), preemption (pausing or removing active work so resources can go elsewhere), or repeated recomputation. A scheduler needs enough headroom, meaning deliberately unused capacity, to avoid memory thrashing—so much eviction and recomputation that useful progress slows down.

Prefill and decode stress hardware differently​

Prompt processing and token decoding also have different shapes. Prefill is the phase that processes the input tokens already present, often with more parallel computation. Decode repeatedly advances active sequences one token step at a time. Workload mix affects which resource becomes the bottleneck.

Prefix reuse can help when requests share common input prefixes, but only when the runtime and workload support it safely. Shared-cache optimization must not cross data-isolation boundaries between users or tasks that should not share state.

Cancellation and completion release cache capacity. A scheduler that keeps abandoned requests alive wastes memory and can reduce capacity for new traffic. Terminal lifecycle and memory lifecycle should be connected.

Measure cache by request class​

The practical service metric is not 'KV cache enabled.' It is how much cache capacity exists, how many active tokens consume it, how frequently requests wait or are preempted, and whether throughput improves as concurrency changes.

Cache accounting should also be tied to request class. A short interactive request and a long document-generation request can consume very different amounts of active-token memory, so one global request-count limit can be misleading. Tracking estimated cached tokens per admitted request makes the controller's decision explainable when capacity becomes tight.

Predict

You keep increasing concurrent active sequences, and throughput eventually falls while cache pressure rises. What is a plausible explanation?

Run the local Lab​

Run:

python3 labs/notebooks/level-14/l14-06-kv-cache.py

The Lab uses a fixed 12,000-token cache budget and admits requests in order.

  1. Run it unchanged. Record the admitted, queued, and free_tokens fields.
  2. In request a, find "cached_tokens": 3000. Before editing, predict whether request b will still fit if only request a grows to 8,000 cached tokens.
  3. Change only request a to "cached_tokens": 8000, then rerun.
  4. Confirm the admission set changes under the same total cache budget: the larger first request leaves too little room for later requests that previously fit.

Loading lab…

Quick Check

1. Why does the KV cache help autoregressive decoding?
2. What generally increases KV-cache pressure?
3. Why keep headroom instead of admitting every request that barely fits?

0 of 3 questions answered.

Explain it back​

Explain how model weights and KV cache compete for device memory. Then describe what happens to admission when average context length doubles.

Key Takeaways

  • KV cache reuses attention state during generation.
  • Active cached tokens consume a concurrency-dependent memory budget.
  • Prefill and decode have different workload shapes.
  • Over-admission can hurt throughput through cache pressure.
  • Completion and cancellation should promptly release request state.

Next Lesson

You have reached the Level 14 checkpoint. Continue to L14.7 — Latency, Throughput, and Cost after reviewing the first six Lessons.

References

Lesson actions

Completion is stored locally on this device.

View progress