Batching and Streaming
Goal
Explain why a serving system groups requests, why waiting for a batch can hurt interactive users, how streaming changes what the user sees, and how scheduling choices trade throughput against delay.
Start with three requests arriving together
Suppose three people ask the same service for different amounts of text:
A: 40 output tokens
B: 200 output tokens
C: 25 output tokens
One simple server could finish A, then B, then C. That is easy to understand, but a GPU can often do more useful work when it processes several sequences together.
Batching groups active requests so one model step can serve more than one sequence. The benefit is better hardware use. The cost is waiting: a request that arrives now may sit in a queue while the server decides what else should run with it.
That trade-off becomes obvious with a fixed batch of 32. During heavy traffic, 32 requests may arrive quickly. Late at night, the first request might wait a long time for 31 more requests that are not coming. A throughput optimization has turned into user-visible delay.
So do not ask only, “What batch size makes the GPU busiest?” Ask two questions together: “How much work does the hardware finish?” and “How long does an individual request wait before useful output begins?”
Continuous batching changes admission
Continuous batching changes the scheduling model. Modern serving runtimes can admit new work as other sequences finish rather than keeping one fixed batch until every sequence completes. vLLM documents continuous batching as one of its throughput-oriented serving features. The exact scheduler changes over time, but the stable idea is dynamic admission of active requests.
Stream for responsiveness
Streaming solves a different problem. Instead of waiting until every output token is generated, a server can send partial output as tokens become available. This can improve time to first token, even though total generation time may remain similar.
Batching and streaming can coexist. The scheduler may batch several active sequences on the device while each client receives its own stream. Request identity is essential so partial outputs are routed to the correct connection.
GPU step 1: [A, B, C]
GPU step 2: [A, B, C]
B finishes → D can enter
client A receives tokens as they are produced
Protect fairness and cancellation
Fairness matters when requests have very different lengths. One extremely long generation should not cause a queue of short interactive requests to wait indefinitely. Schedulers can use policies such as admission limits, age, request class, or token budget to balance throughput and responsiveness.
Cancellation also interacts with streaming. A client may disconnect after receiving enough output. The service should stop scheduling useless future work when safe to do so, release request-specific KV cache, and record a cancellation terminal reason.
Measure queue time separately
Measure queue time separately from compute time. A faster model does not help much if requests spend most of their latency waiting for admission. Likewise, a larger batch that improves tokens per second may hurt interactive time to first token.
The practical goal is not maximum batch size. It is a scheduling policy that matches workload goals: interactive chat may prioritize first-token latency, while offline batch generation may prioritize total throughput and cost.
Predict
Run the local Lab
Run:
python3 labs/notebooks/level-14/l14-03-batching-streaming.py
The Lab simulates fixed arrivals under a maximum batch size and wait deadline.
- Run it unchanged with
wait_deadline = 2. Record the batch count and longest queue wait. - Before editing, predict the tradeoff if the scheduler may wait up to 10 ms: should requests form fewer/larger batches, and should the oldest request be allowed to wait longer?
- Change only
wait_deadline = 2towait_deadline = 10, then rerun. - Confirm the same arrivals now form fewer batches while the longest permitted queue wait increases. The workload did not change; only the batching deadline did.
Loading lab…
Write the core logic yourself
Open:
labs/notebooks/level-14/l14-03-batching-streaming-exercise.py
Implement bounded batch formation using both maximum batch size and wait deadline. The supplied arrivals must split into two batches for the correct reason, not because of a hard-coded answer.
Run:
python3 labs/notebooks/level-14/l14-03-batching-streaming-exercise.py
The starter intentionally stops at TODO until you implement the missing logic. A correct solution reaches the final PASS: marker. Use the solved deterministic Lab as a comparison only after your own attempt.
Quick Check
Explain it back
Compare an interactive chat workload with an offline document-generation workload. Explain how you would choose batch-wait time, maximum concurrency, and streaming behavior for each.
Key Takeaways
- Batching trades waiting time for hardware efficiency.
- Continuous batching dynamically mixes active requests.
- Streaming targets time to first output rather than total compute alone.
- Fairness and cancellation are scheduler responsibilities.
- Measure queue and compute latency separately.
Next Lesson
Next, L14.4 — Model Loading and Memory accounts for where serving memory is actually spent.
References
- vLLM, Documentation.
Completion is stored locally on this device.