본문으로 건너뛰기

Checkpoint — Inference Service Mechanics

Before moving into GPU scheduling and production operations, make sure the first half of Level 14 forms one consistent resource model.

A service is more than a generation function. You should be able to separate startup/model lifecycle from request handling, distinguish liveness from readiness, validate bounded inference inputs, and attach deployment identity to responses and traces.

You should also be able to explain the relationship between batching and streaming. Batching groups work to improve accelerator efficiency. Streaming lets a client see partial output before generation is complete. They solve different latency problems and can be used together.

Memory is the next boundary. Model weights are only one pool. Runtime workspace and request-dependent KV cache also consume device memory. As active sequence length and concurrency grow, cache pressure can become the limiting resource even when the weights fit comfortably.

Quantization can reduce weight memory, but it is not a guaranteed proportional speedup. Hardware support, serving runtime, kernels, quality, and workload shape all matter. Treat a quantized deployment as a separately measured candidate.

Finally, explain why the KV cache improves autoregressive generation and why over-admission can still hurt throughput. The scheduler needs a safe cache budget and headroom, not merely the largest possible count of active requests.

Self-check​

You are ready to continue if you can answer:

  1. What belongs in startup versus a request handler?
  2. Why are liveness and readiness different?
  3. How do batching and streaming affect different latency components?
  4. Which serving-memory pools exist beyond weights?
  5. Why can quantization reduce memory without producing the same percentage latency improvement?
  6. What makes KV-cache demand grow?

If an answer is vague, rerun that Lesson's deterministic Lab and explain the changed output before continuing.

Check your reasoning with an applied serving case

Suppose a service is alive, but the model failed to load. Its liveness check can still say the process is running, while readiness must fail so traffic is not sent to it.

Now suppose the model loads, but long prompts and high concurrency fill the KV-cache budget. The weights still fit, yet admission may need to slow or reject new work. That is why serving memory includes more than weights.

Use this case to check the other ideas:

  • startup loads and verifies shared artifacts; request handlers validate bounded per-request inputs;
  • batching groups work for better device use, while streaming exposes partial output sooner;
  • quantization can reduce weight memory, but latency depends on kernels, hardware, runtime, and workload, so measure it rather than assuming the same percentage speedup;
  • KV-cache demand grows with active sequence length and concurrency, so the scheduler needs headroom instead of admitting the largest possible request count.

Next Lesson

Continue with L14.7 — Latency, Throughput, and Cost.

Lesson actions

Completion is stored locally on this device.

View progress