Level 0 project
Data Detective: Build and Explain a Tiny Predictor
Start from the canonical project files, validate the result, and keep the evidence you need to explain what you built.
Start here
Launch lesson: Debug and Improve a Tiny Predictor
No prerequisite Project.
Goal
Demonstrate that you can define a tiny prediction task, preserve fair evaluation evidence, compare a baseline, investigate a failure, and explain what the evidence does and does not support.
Task
Path A — Core: browser evidence + explanation
Use this path if you have not learned repository, terminal, or Python implementation workflow yet.
- Open Lesson 0.12 and run
lab-l00-12unchanged. - Record the starter
debug_record, especiallybefore,after, andbaseline. - Find
candidate_threshold = 3in the Lab and change only3to2. - Before running, predict what will happen to input
2. - Run again and record the new
before,after, andbaselinevalues. - Complete
CORE_REPORT_TEMPLATE.mdfrom this folder, or copy the same prompts from the learner-facing Project Workbench page.
This path is complete when your evidence report demonstrates all five rubric areas. You do not need to implement Python functions or run a terminal validator for the Core path.
Path B — Builder: local Python implementation
Use this path when you want to practice turning the same reasoning into code.
If repository/terminal workflow is new, read projects/PROJECT_WORKBENCH.md first. It defines repository root, terminal, validator, and the setup-debugging order.
Setup
No third-party packages are required. Use Python 3.11+ from the repository root.
Open projects/starters/l00/data_detective.py and complete the marked TODOs.
Your program must:
- treat the numeric measurement as the feature and
readyas the label; - choose a threshold using only
TRAIN_EXAMPLES; - evaluate the chosen threshold on
TEST_EXAMPLESwithout tuning on the test answers; - compare test mistakes with a majority-label baseline learned from the training labels;
- produce a short report containing the chosen threshold, model mistakes, baseline mistakes, and a cautious conclusion;
- investigate the intentional
NOISY_TEST_EXAMPLESscenario and write a short debug note.
Builder deliverables
- completed
data_detective.py; - terminal output from the validation command;
- a short
debug-note.mdcontaining: observed failure, hypothesis, one change, result, and next step; - a 4–8 sentence explanation of what the held-out evidence supports and one limitation.
Builder validation
From the repository root, run:
python projects/tests/l00/validate_submission.py projects/starters/l00/data_detective.pyExpected success evidence ends with:
PASS: p00-data-detective objective checksThe checks are behavioral. Your printed wording does not need to match a reference solution exactly.
Validation
Run these commands from the downloaded Project folder or the public materials repository root.
python projects/tests/l00/validate_submission.py projects/starters/l00/data_detective.pyRubric
This rubric maps directly to the Level 0 exit skills. Core and Builder are two evidence paths to the same learning outcome. A learner must not lose Level 0 credit merely because local Python or terminal workflow has not been learned yet.
Accepted evidence paths:
- Core browser path: canonical Lesson 0.12 Lab evidence + completed Core Evidence Report;
- Builder local path: completed
data_detective.py+ validator evidence + debug/explanation deliverables.
Judge the reasoning and reproducibility appropriate to the chosen path. Do not award extra conceptual credit merely for using more code.
1. Features and labels — 20 points
- 18–20: Correctly identifies the numeric input as the feature and the Boolean target as the label; explanation clearly distinguishes prediction from label.
- 12–17: Reasoning is mostly correct but explanation is incomplete or mixes one term.
- 1–11: Feature/label roles are confused in evidence or prose.
- 0: No usable evidence.
2. Train/test separation and fair evaluation — 20 points
- 18–20: Explains that the rule is chosen without using final held-out answers; held-out examples are used for evaluation; explains why repeated tuning on final test answers weakens independence.
- 12–17: Correct boundary with weak explanation, or one minor evaluation mistake that is identified and corrected.
- 1–11: Test answers guide model selection without recognizing the problem, or the split is not reproducible.
- 0: No train/test distinction.
3. Predictor, baseline, and objective evidence — 20 points
- 18–20 Core: Correctly records the canonical Lab's before/after/baseline evidence, traces the changed threshold to the changed prediction, and interprets the baseline fairly.
- 18–20 Builder: Objective validator passes; model and majority baseline are evaluated on the same held-out examples; result is interpreted correctly.
- 12–17: Predictor evidence and baseline are present with one small defect or incomplete interpretation.
- 1–11: A predictor result exists but the comparison is unfair, untraceable, or missing a baseline.
- 0: No functioning or inspectable predictor evidence.
4. Failure analysis and debugging — 20 points
- 18–20: Learner identifies a specific failure, states a plausible hypothesis, changes or checks one thing at a time, records evidence, and states a next step. The Core path may use the deliberate threshold
3 → 2unsuccessful change; the Builder path may useNOISY_TEST_EXAMPLES. - 12–17: Failure is found and partially explained, but cause/evidence/next-step chain is incomplete.
- 1–11: Failure is hidden, fixed by unrelated changes, or discussed without evidence.
- 0: No debug attempt.
5. Reproducibility and communication — 20 points
- 18–20 Core: Records the canonical Lesson/Lab, exact one-line edit, before/after/baseline values, observation versus conclusion, and one limitation so another learner can repeat the browser experiment.
- 18–20 Builder: Provides validation command/output and enough settings/evidence to reproduce the local run; distinguishes observation from conclusion and states one limitation.
- 12–17: Mostly reproducible evidence with an incomplete record or limitation statement.
- 1–11: Result depends on undocumented changes or explanation makes claims broader than the evidence.
- 0: Submission cannot be reproduced or explained.
Performance bands
- 90–100: Ready to advance; all Level 0 exit skills demonstrated.
- 75–89: Meets core outcome with one or two specific areas to strengthen.
- 60–74: Partial mastery; repeat the relevant Lesson/Lab before advancing.
- 0–59: Major Level 0 skills are missing; revise with rubric evidence.
A passing local validator is required only for the Builder local path. It is not a prerequisite for full Level 0 conceptual mastery on the Core browser path.