본문으로 건너뛰기

Level 7 project

Reliable LLM Workflow

Start from the canonical project files, validate the result, and keep the evidence you need to explain what you built.

Launch lesson: Reliable LLM Workflow Review

Prerequisite project: Train and Ship a Tiny LLM

Goal

Build and review a model-facing workflow whose reliability comes from explicit contracts rather than one lucky generated answer.

Your workflow must keep these boundaries visible:

task contract
→ context selection
→ model-facing request
→ decoding configuration
→ structured output
→ schema/provenance validation
→ abstention or accepted answer
→ fixed-case evaluation

Task

Complete the TODOs in workflow.py:

  1. word_overlap(query, text)
  2. retrieve(query, documents, k=1)
  3. validate_model_output(output, supplied_source_ids)
  4. evaluate_cases(cases)
  5. validate_run(data)

Produce a workflow-run.json that records:

  • workflow, model, and prompt identity;
  • decoding/seed policy;
  • source catalog;
  • fixed evaluation cases;
  • the exact source IDs supplied to each case;
  • structured output for each case;
  • supported/unsupported expectations;
  • one failure/debug record;
  • limitations.

Validation

Run these commands from the downloaded Project folder or the public materials repository root.

python projects/tests/l07/validate_submission.py \
projects/starters/l07/workflow.py \
path/to/your/workflow-run.json

Rubric

Score each criterion from 0–4. A submission should score at least 16/20 overall and must not score 0 on provenance/validation or evaluation.

Criterion4 — Strong evidence3 — Meets2 — Partial1 — Weak0 — Missing
Task and prompt interfaceTask, input boundary, constraints, output contract, and abstention behavior are explicit and testableClear interface with minor omissionsInterface exists but some requirements remain vagueMostly prompt prose without a checkable contractNo coherent task contract
Context and provenanceContext-selection strategy is documented; every accepted answer cites a source supplied for that exact request; untrusted content remains visibly untrustedCorrect source/context handling with small gapsProvenance exists but is incomplete or inconsistentSource catalog exists without per-request groundingNo source/context evidence
Structured validation and trust boundariesParse/schema/provenance/business rules are separated; unsupported output abstains; sensitive actions are bounded outside the modelCorrect validation flow with minor gapsSome validation but important boundary missingTrust placed mainly in prompt wordingNo meaningful validation
Evaluation and failure analysisFixed diverse cases, aggregate + failed-case evidence, controlled conditions, and one failure→evidence→hypothesis→fix→result traceComplete evaluation with minor gapsEvaluation present but shallow or overfitA few cherry-picked samples onlyNo evaluation
Reproducibility and communicationPrompt/model/evaluator/context/decoding identities, commands, limitations, and optional live-model differences are recorded clearlyReproducible package with small omissionsImportant identity/configuration missingMostly screenshots/output dumpNo reproducible package

Exit-skill mapping

  • Design/evaluate prompts: Task and prompt interface; evaluation.
  • Manage context: Context and provenance.
  • Request structured outputs: Structured validation.
  • Recognize hallucination/injection risks: Trust boundaries; failure analysis.
  • Measure reliability: Evaluation and reproducibility.

Objective checks

The reduced validator checks retrieval behavior, structured-output provenance, abstention, fixed-case evaluation, run metadata, and the intentional unsupplied-source failure.

Human review scores whether the actual prompt/context choices, trust boundaries, evaluation coverage, and conclusions are technically convincing.

Assessment guardrail

A live provider model is optional for repository acceptance. If one is used, model output does not need to match the reference text. The project rewards explicit contracts and evidence rather than provider-specific prompt tricks.

← Back to all projects