Skip to main content
L14.7

Latency, Throughput, and Cost

Goal

Measure time to first token, end-to-end latency, throughput, utilization, and cost together so optimization targets the real service objective.

One speed number is not enough​

Consider two chat services. Service A starts showing words after 0.4 seconds and finishes after 5 seconds. Service B shows nothing for 3 seconds but finishes after 4 seconds. B completes sooner, yet A may feel more responsive because the user sees useful progress much earlier. One number cannot describe both experiences.

A production service therefore needs more than one performance number. End-to-end latency tells how long the client waited for completion. Interactive generation also cares about time to first token (TTFT): the delay before the first streamed token reaches the client. Throughput asks a different question—how much useful work the system completes per unit of time.

These measurements can move in opposite directions. Waiting briefly to form a larger batch may increase total tokens per second while making the first token slower for an individual chat request. The goal is not to make every metric as large or small as possible; it is to choose tradeoffs that match the workload.

service A: TTFT 0.4 s, end-to-end 5.0 s
service B: TTFT 3.0 s, end-to-end 4.0 s

The faster completion and the faster first response are different properties.

Measure throughput and latency separately​

Throughput measures completed work over time. Depending on the service, useful units include requests per second, input tokens per second, output tokens per second, or completed jobs per minute. Choose a unit that matches the workload rather than quoting one number without context.

Queue wait, prefill, decode, network transfer, and client streaming all contribute to latency. Trace those components separately. A slow request may come from an overloaded queue even when GPU execution is healthy.

Use tail latency and unit cost​

Percentiles matter because averages hide tail behavior. p50 describes a typical request, while p95 or p99 reveals slower users. Tail latency often grows first when capacity is close to saturation.

Cost should be tied to useful work. If a configuration doubles throughput on one GPU but also doubles total GPU count, unit cost may not improve. Track resource-hours, request/tokens processed, and any external service charges needed for the workload.

Go deeper: compute p95 and unit cost

For a tiny latency sample, sort 20 TTFT measurements:

180, 190, 200, 205, 210,
215, 220, 225, 230, 235,
240, 245, 250, 260, 270,
280, 300, 340, 480, 900 ms

Using the simple nearest-rank rule:

rank = ceil(0.95 × 20) = 19
p95 = 19th sorted value = 480 ms

The 900 ms request still matters, but p95 is not the maximum. Production libraries may use interpolation or another percentile convention, so record the method when exact reproduction matters.

Now normalize cost. Suppose a test uses two GPUs priced at $2.40/hour each for 30 minutes and completes 12,000 requests:

GPU cost = 2 × $2.40/hour × 0.5 hour
= $2.40

cost per request
= $2.40 / 12,000
= $0.0002

cost per 1,000 requests
= $0.20

That is compute-only cost for this example. A real unit-cost calculation may also include storage, network, external APIs, idle reserve, and human operations.

Watch tradeoffs move the bottleneck​

Optimization can move the bottleneck. Quantization may reduce weight memory enough to admit more requests, which then increases KV-cache pressure. Larger batches may improve device utilization while increasing queue delay. The correct scorecard keeps latency, throughput, quality, errors, and cost visible together.

Service-level objectives turn measurements into operating targets. For example, a workload might require p95 time-to-first-token below a threshold, task error rate below a threshold, and no overload-induced crashes. The exact numbers are workload-specific.

Connect metrics to traces​

OpenTelemetry can carry traces and metrics that connect these components. The important design is consistent request/deployment identifiers so aggregate metrics can be drilled into individual slow traces.

Capacity experiments should change one control at a time when possible. Run the same fixed workload at different concurrency levels and record where throughput stops scaling and tail latency accelerates. That knee is more informative than one maximum-throughput benchmark.

Predict

A configuration increases total tokens per second but doubles p95 time to first token for chat users. Is it automatically better?

Run the Docker-environment Lab preflight​

This activity is registered for the Docker-oriented production environment used in the second half of Level 14. Start with the deterministic Python preflight so service-logic failures remain distinguishable from container, GPU, cluster, or credential setup problems.

Run:

python3 labs/notebooks/level-14/l14-07-slo-metrics.py

The Lab holds the request workload fixed and changes only serving concurrency.

  1. Run it unchanged with CONCURRENCY = 4. Record p95 TTFT, p95 end-to-end latency, tokens per second, and cost per 1k tokens.
  2. Before editing, predict the tradeoff if concurrency doubles: throughput can improve because more requests share each wave, but per-request TTFT and token time can worsen from contention.
  3. Change only CONCURRENCY = 4 to CONCURRENCY = 8, then rerun.
  4. Compare all four metrics. Explain why a throughput gain does not make the configuration automatically better if tail latency moves in the wrong direction.

Loading lab…

Quick Check

1. What does time to first token measure?
2. Why examine p95 instead of only average latency?
3. How should unit cost be interpreted?

0 of 3 questions answered.

Explain it back​

Design a serving scorecard with TTFT, end-to-end p95, tokens/sec, error rate, and cost per useful unit. Explain one tradeoff you would refuse despite higher throughput.

Key Takeaways

  • Latency has several components and user-facing meanings.
  • Throughput needs workload-appropriate units.
  • Tail percentiles reveal saturation earlier than averages.
  • Cost should be normalized to useful work.
  • Optimize against explicit service objectives, not one benchmark.

Next Lesson

Next, L14.8 — GPU Scheduling Basics models devices as scarce workers with memory and compute constraints.

References

Lesson actions

Completion is stored locally on this device.

View progress