Rollouts, Canaries, and Rollbacks
Goal
Release a new service or model gradually, compare it with clear gates, and roll back when critical quality or reliability rules fail.
Change a little traffic first
Imagine a school cafeteria changing a recipe. Serving the new recipe to every student at once makes a mistake expensive. A safer test is to use the new recipe at one counter, keep the old recipe at the others, and compare clear results before expanding it.
A production AI change has the same problem even when the API looks unchanged. A new model, quantization format, runtime, batch rule, or container can change quality, memory, latency, or errors.
A rollout controls how much traffic sees the new version. A gradual rollout sends only part of the traffic first so you can compare the candidate with the known version, called the baseline.
Define canary evidence and release gates
A canary is that small early release. For example, 5% of eligible traffic might use the candidate while the rest stays on the baseline. Record which requests used which version so their metrics stay separate. The percentage is a deployment choice, not a universal rule.
Write release gates before seeing candidate results. Examples are zero critical authorization failures, a maximum error rate, a minimum task-success rate, a p95 latency limit, and a minimum memory reserve. Predefined gates make it harder to excuse a bad result after the fact.
candidate traffic: 5%
critical authorization failures: 0 required
p95 latency: must stay below the declared limit
if a gate fails → route future traffic back to baseline
Rollback is not undo
Kubernetes Deployments are one example of gradual replacement and rollback. The broader rule is simple: releases should be observable and reversible.
A rollback changes future deployment and routing. It does not undo every side effect the candidate already caused. External writes need their own idempotency and reconciliation rules.
Compare representative steady-state traffic
Warm-up can distort comparisons. A new replica may still be loading, compiling, or filling caches. Do not use a few cold requests as steady-state evidence unless startup time is the thing you are measuring.
The canary traffic should also resemble the traffic you care about. If the first 1% is unusual, its result may not predict a full rollout.
A safe release combines pre-release tests, exact candidate identity, limited exposure, comparable telemetry, explicit gates, and a tested rollback action.
Predict
Run the Docker-environment Lab preflight
This activity is registered for the Docker-oriented production environment used in the second half of Level 14. Start with the deterministic Python preflight so service-logic failures remain distinguishable from container, GPU, cluster, or credential setup problems.
Run:
python3 labs/notebooks/level-14/l14-11-rollout-gate.py
The Lab promotes a canary only when every predefined gate passes.
- Run it unchanged. All three gates should be
okand the decision should bepromote canary. - In
canary, find"critical_violations": 0. Before editing, predict the release decision if exactly one critical violation is observed while latency and error rate stay unchanged. - Change only
"critical_violations": 0to1, then rerun. - Confirm the critical gate becomes
FAIL, the decision changes toroll back canary, and the script exits non-zero.
Loading lab…
Write the core logic yourself
Open:
labs/notebooks/level-14/l14-11-rollout-gate-exercise.py
Implement the canary release gate from predefined error-rate, p95-latency, and critical-violation limits. A single critical violation must block release even when the average metrics look good.
Run:
python3 labs/notebooks/level-14/l14-11-rollout-gate-exercise.py
The starter intentionally stops at TODO until you implement the missing logic. A correct solution reaches the final PASS: marker. Use the solved deterministic Lab as a comparison only after your own attempt.
Quick Check
Explain it back
Design a canary rollout with traffic percentage, warm-up handling, five release gates, and one rollback trigger. Explain which previous effects cannot be undone automatically.
Key Takeaways
- Serving changes need gradual, observable release control.
- Canaries limit exposure while collecting real evidence.
- Define release gates before evaluating the candidate.
- Rollback changes deployment state, not all past effects.
- Keep candidate and baseline telemetry distinguishable.
Next Lesson
Next, L14.12 — Capacity Planning and Failure Recovery connects workload demand to replicas, queues, headroom, and degraded operation.
References
- Kubernetes, Deployments.
Completion is stored locally on this device.