Skip to main content

Level 14 project

Production AI Service

Start from the canonical project files, validate the result, and keep the evidence you need to explain what you built.

Launch lesson: Production AI Operations Workshop

Prerequisite project: Interoperable Agent System

Goal

Build the operational controller around a production inference service.

Task

Complete the TODO functions in service.py:

  1. bounded inference request validation;
  2. deadline/size-bounded batching;
  3. memory admission with reserve headroom;
  4. immutable deployment identity;
  5. serving metric aggregation;
  6. rollout release gates;
  7. replica capacity planning;
  8. raw-evidence run validation.

The required acceptance path is deterministic and does not require a GPU, live model server, Kubernetes cluster, cloud account, network call, or secret.

Run:

python3 projects/tests/l14/validate_submission.py \
  projects/starters/l14/service.py \
  projects/tests/l14/fixtures/passing/service-run.json

Reference acceptance:

python3 projects/tests/l14/test_reference.py

Level Labs:

python3 labs/notebooks/level-14/test_labs.py

Validation

Run these commands from the downloaded Project folder or the public materials repository root.

python3 projects/tests/l14/validate_submission.py \

Rubric

Total: 100 points.

  • API and admission control — 20: request bounds and resource admission are deterministic and reject unsafe work before expensive execution.
  • Batching and resource accounting — 15: batch formation and memory reserve rules remain explicit and testable.
  • Deployment reproducibility — 15: image, service, model, runtime, and configuration identity are preserved.
  • Observability and performance — 15: success and latency evidence can be recomputed from raw run records.
  • Rollout and recovery — 20: critical violations block release, deployment mismatch is visible, and capacity logic preserves operational headroom.
  • Reproducibility — 15: offline tests and the Docker bridge separate controller correctness from GPU/cloud environment setup.

Full credit requires recomputing critical release evidence rather than trusting fixture fields such as admitted, healthy, or release_passed when primary evidence is available.

← Back to all projects