본문으로 건너뛰기
L14.13

Production AI Operations Workshop

Goal

Combine request validation, batching, memory limits, deployment identity, telemetry, rollout gates, and recovery in one testable service controller.

Follow one request across the service​

Follow one request before looking at the whole controller. A client sends "summarize this text" with max_tokens=200. The API first checks that the fields and limits are valid. Admission then asks whether the current service has enough safe capacity. The scheduler decides when and where the request runs. The model performs inference. The service streams or returns the output. Telemetry records what happened under the exact deployment version.

A request can fail at any one of those boundaries for a different reason. Invalid input should fail before expensive work starts. A valid request may still wait because memory is full. A healthy inference can still belong to a bad rollout if the candidate violates a release gate. Keeping the boundaries separate makes the earliest cause visible.

The same request can be written as a concrete sequence:

  1. API validation
  2. admission
  3. scheduling
  4. model execution
  5. streaming or final response
  6. telemetry
  7. terminal accounting

Pin deployment identity and admission​

The deployment record names the exact service image, model revision, serving runtime, and configuration. Carry that identity with every fixture so measurements from two candidates are not mixed.

{"request_id":"r-17","deployment":"v18","admitted":true,"queue_ms":42,"ttft_ms":310,"terminal":"ok"}

Admission checks both request limits and available memory. A valid request can still wait or receive an overload response when the current replica cannot fit it safely. Clean rejection is better than crashing the shared model process.

Schedule and account for memory​

The scheduler creates batches under a wait deadline and concurrency limit. Record queue time, time to first token, and full generation time separately. Then a throughput improvement cannot hide a large increase in interactive wait time.

Memory accounting includes weights, KV-cache budget, workspace, and reserve. The workshop does not emulate a real GPU kernel. It checks whether controller decisions stay inside the declared resource budget.

Use telemetry and rollout gates​

Telemetry records the deployment version, request class, latency pieces, error category, and capacity signals. Keep request IDs in trace-like records instead of aggregate metric labels.

The rollout fixture compares a baseline with a candidate. The candidate passes only when normal thresholds and critical zero-tolerance rules all hold. If a gate fails, the controller chooses the known rollback deployment.

Test recovery and recompute evidence​

Recovery tests also remove one replica. The remaining system must respect safe capacity, queue limits, and degraded-mode rules. A test that works only when every resource is healthy is not enough.

The project validator recomputes critical metrics from raw requests, deployments, and events. It does not trust fields such as admitted, release_passed, or healthy when those values can be derived from primary evidence.

When something fails, debug the earliest broken boundary. An out-of-memory crash may begin with bad admission accounting. A latency regression may begin in the queue. A rollout comparison may be wrong because telemetry mixed two versions. Clear evidence at each boundary makes the cause easier to find.

Predict

A candidate serving version has good average latency but its raw events show a memory-admission violation. The release rule allows zero such violations. What should the workshop decide?

Run the Docker-environment Lab preflight​

This activity is registered for the Docker-oriented production environment used in the second half of Level 14. Start with the deterministic Python preflight so service-logic failures remain distinguishable from container, GPU, cluster, or credential setup problems.

Run the passing production-service fixture first:

python3 labs/notebooks/level-14/l14-13-integration-check.py projects/reference/l14/sample-run/service-run.json

Confirm the validator recomputes zero memory-admission violations and a passing release.

For a one-variable experiment, copy the passing fixture:

cp projects/reference/l14/sample-run/service-run.json /tmp/l14-service-run.json
python3 labs/notebooks/level-14/l14-13-integration-check.py /tmp/l14-service-run.json

In run a of the copied JSON, find "estimated_memory_gb": 4 while "available_memory_gb": 8. Before editing, predict which recomputed metric and release gate should change if the admitted request needs 10 GB instead. Change only that one estimated_memory_gb value to 10, then rerun the validator on /tmp/l14-service-run.json.

The validator should now count one memory-admission violation and reject the release while request validity, deployment identity, and the other unchanged evidence remain separate.

Finally, run the repository's intentional multi-failure fixture:

python3 labs/notebooks/level-14/l14-13-integration-check.py projects/reference/l14/failure-run/service-run.json

Use it as a debugging challenge. Trace each release-blocking metric back to the raw request that caused it. Do not trust the stored summary by itself.

Loading lab…

Quick Check

1. Which check belongs in request admission rather than startup readiness or client preference?
2. Why carry deployment identity into telemetry?
3. How should the project validator treat derived release flags?

0 of 3 questions answered.

Explain it back​

Walk through one production request from API validation to deployment telemetry, then explain how the same evidence enters a canary release and capacity-recovery decision.

Key Takeaways

  • Production inference needs explicit service and resource boundaries.
  • Deployment identity must remain attached to metrics and traces.
  • Admission protects shared model/runtime capacity.
  • Rollout gates combine performance, quality, and critical safety/reliability rules.
  • Project validation should recompute derived operational metrics from raw evidence.

Next Lesson

You are ready for the Level 14 project: build and validate the Production AI Service.

Bridge to Level 15: Level 14 teaches how to keep an AI service available, measurable, and recoverable. Running well is not the same as being safe or ready to release. Level 15 reuses the evidence your service now produces: tests, traces, metrics, versions, and failure records. That evidence supports quality evaluation, security testing, and the final release decision.

References

Lesson actions

Completion is stored locally on this device.

Level project unlocked: Production AI Service

View progress