본문으로 건너뛰기

Level 14 project

Production AI Service

Start from the canonical project files, validate the result, and keep the evidence you need to explain what you built.

Launch lesson: Production AI Operations Workshop

Prerequisite project: Interoperable Agent System

Goal

Build the operational controller around a production inference service.

Task

Complete the TODO functions in service.py:

  1. bounded inference request validation;
  2. deadline/size-bounded batching;
  3. memory admission with reserve headroom;
  4. immutable deployment identity;
  5. serving metric aggregation;
  6. rollout release gates;
  7. replica capacity planning;
  8. raw-evidence run validation.

The required acceptance path is deterministic and does not require a GPU, live model server, Kubernetes cluster, cloud account, network call, or secret.

Run:

python3 projects/tests/l14/validate_submission.py \
  projects/starters/l14/service.py \
  projects/tests/l14/fixtures/passing/service-run.json

Reference acceptance:

python3 projects/tests/l14/test_reference.py

Level Labs:

python3 labs/notebooks/level-14/test_labs.py

Validation

Run these commands from the downloaded Project folder or the public materials repository root.

python3 projects/tests/l14/validate_submission.py \

Rubric

Total: 100 points.

  • API and admission control — 20: request bounds and resource admission are deterministic and reject unsafe work before expensive execution.
  • Batching and resource accounting — 15: batch formation and memory reserve rules remain explicit and testable.
  • Deployment reproducibility — 15: image, service, model, runtime, and configuration identity are preserved.
  • Observability and performance — 15: success and latency evidence can be recomputed from raw run records.
  • Rollout and recovery — 20: critical violations block release, deployment mismatch is visible, and capacity logic preserves operational headroom.
  • Reproducibility — 15: offline tests and the Docker bridge separate controller correctness from GPU/cloud environment setup.

Full credit requires recomputing critical release evidence rather than trusting fixture fields such as admitted, healthy, or release_passed when primary evidence is available.

← Back to all projects