Skip to main content
L14.4

Model Loading and Memory

Goal

Estimate serving memory across model weights, KV cache, runtime workspace, and request state so admission decisions use explicit budgets rather than guesswork.

Budget the whole GPU, not just the weights​

Picture a desk with limited surface area. A large textbook stays open all day. Each student request adds sticky notes that must remain available while that request is active. The program also needs scratch paper for temporary calculations, and you deliberately leave part of the desk empty so a small surprise does not knock everything onto the floor.

GPU memory works in a similar way. Model weights are the large, mostly fixed textbook. The KV cache and other request state are sticky notes that grow with active sequences. Runtime workspace is scratch space used while operations run. A safety reserve is memory you choose not to fill during normal admission.

For a simplified example, imagine a 24 GB device where weights use 14 GB, runtime workspace uses 2 GB, and you reserve 2 GB for safety. That leaves about 6 GB for request-dependent state. If one long active request is estimated to need 1.5 GB of that state, four such requests fit the simplified budget but a fifth would cross it. Real runtimes have more details, but the reasoning habit is the same: add the memory pools before deciding what can start.

24 GB total
- 14 GB weights
- 2 GB workspace
- 2 GB safety reserve
= 6 GB left for request-dependent state

Separate fixed and request-dependent memory​

Model weights are the large, mostly fixed part. A rough estimate is parameter count × bytes per stored value. Real systems also need some extra space for metadata, sharding, and runtime buffers.

Generation adds the KV cache. It stores attention keys and values for tokens that were already processed. Longer prompts, more generated tokens, and more active requests make this cache grow. The cache saves compute, but it competes with the other memory pools.

Runtime code may also need temporary workspace. Some systems reserve memory in pools. Moving data between host memory and GPU memory can be expensive, so placement matters.

Free memory is not guaranteed capacity​

A "free memory" number can also be misleading. Memory may be reserved or split into pieces that are hard to reuse for one large allocation. Serving runtimes often manage blocks or pools to make reuse more predictable.

Admission control decides whether a new request may start now. If the request could push memory above the safe budget, the service should queue or reject it. That is safer than letting one request crash the shared model process.

Plan loading and offloading​

Model loading time also affects operations. A replacement process may take seconds or minutes before it becomes ready, depending on model size, storage, compilation, and device initialization. Rollout plans should account for warm-up and not assume a new replica is ready immediately after process start.

Offloading moves some weights or cache between GPU and host memory. It can reduce GPU pressure, but transfers add latency and use bandwidth. It is a tradeoff, not free capacity.

Ask four concrete questions: Which memory pools exist? Which are mostly fixed? Which grow with active requests or sequence length? How much reserve keeps the process healthy? Those answers are more useful than one headline number for "GPU memory used."

A simple budget makes this concrete. Suppose the weights consume most of a GPU but leave several gigabytes free. That remaining space is not automatically safe concurrency. Active sequences still need KV cache and temporary runtime memory, and the process needs reserve for variation. If ten requests fit but the eleventh crosses the safe limit, admission control should queue or reject the eleventh instead of treating the remaining bytes as guaranteed capacity.

Predict

A model's weights fit in GPU memory, but adding many long generation requests causes out-of-memory failures. What likely grew?

Run the local Lab​

Run:

python3 labs/notebooks/level-14/l14-04-model-memory.py

The Lab separates fixed model memory from request-dependent KV-cache memory.

  1. Run it unchanged with active_sequences = 8. Record parts_gb and the remaining headroom.
  2. Before editing, predict which memory component changes if the active sequence count doubles. Weights, workspace, and reserve must stay fixed.
  3. Change only active_sequences = 8 to active_sequences = 16, then rerun.
  4. Confirm only kv_cache doubles. The total now exceeds the 24 GB device budget, so the script reports that the workload does not fit.

Loading lab…

Quick Check

1. Why is model-file size an incomplete serving-memory estimate?
2. What should admission control do near a safe memory limit?
3. What tradeoff can memory offloading introduce?

0 of 3 questions answered.

Explain it back​

Create a memory budget with weights, KV cache, workspace, and reserve. Explain which parts are mostly fixed and which grow with active requests.

Key Takeaways

  • Serving memory includes more than model weights.
  • KV cache grows with active generation state.
  • Admission should use a safe memory budget.
  • Model loading and warm-up affect rollout readiness.
  • Offloading trades device capacity for transfer cost.

Next Lesson

Next, L14.5 — Quantization for Serving examines how lower-precision representations change memory, compatibility, and quality.

References

Lesson actions

Completion is stored locally on this device.

View progress