Debug and Improve a Tiny Predictor
Goal
By the end of this lesson, you can use a repeatable debugging loop to inspect one failure, form a hypothesis, change one thing, compare the result fairly, and decide whether the evidence supports the change.
Debugging is not random tweaking
When a prediction is wrong, changing numbers until the score improves can work by luck. It does not tell you why the change helped.
Use this loop instead:
- Find one failure.
- Describe the evidence.
- Propose one possible cause.
- Make one change that tests that cause.
- Evaluate under the same conditions.
- Compare with the original setting and a simple baseline.
- Explain what the result supports.
Start from one visible failure
Training examples:
| Input | Label |
|---|---|
| 1 | False |
| 2 | False |
| 3 | True |
| 4 | True |
| 5 | True |
Held-out examples:
| Input | Label |
|---|---|
| 2 | False |
| 3 | True |
| 4 | True |
| 6 | True |
The initial predictor uses threshold 4:
predict
Truewhen input >= 4
On the held-out set, input 3 is wrong: the predictor says False, but the label is True.
So the initial predictor makes 1 mistake.
A reasonable hypothesis is:
The threshold may be one step too high.
That suggests one exact change: threshold 4 → 3.
With threshold 3, inputs 2, 3, 4, and 6 are all predicted correctly, so we expect the mistake count to fall from 1 to 0.
Compare with a baseline
The training labels contain more True values than False values. A majority baseline therefore predicts True for every input.
On the held-out set, that baseline misses input 2, so it also makes 1 mistake.
Now we have context:
- initial threshold
4: 1 mistake; - candidate threshold
3: expected 0 mistakes; - majority baseline: 1 mistake.
Run the debugging record
- Click Run without changing anything.
- Read the printed
debug_recorddictionary. - Find
failure,hypothesis,change,before,after, andbaseline. - Confirm that the starter reports
before: 1,after: 0, andbaseline: 1.
Loading lab…
The evidence supports a careful statement:
On this held-out set, lowering the threshold from 4 to 3 removed the observed mistake and beat the majority baseline.
It does not prove that threshold 3 is best for every future example.
Test a hypothesis that does not help
Now make one exact edit in the Lab:
candidate_threshold = 3
Change only 3 to 2.
Before clicking Run, predict what will happen to held-out input 2.
With threshold 2, input 2 becomes True, but its label is False. The candidate therefore makes 1 mistake again.
Run the Lab and compare before, after, and baseline.
This result is useful because it rejects the broad idea that any lower threshold must improve the model.
A useful record is:
- Failure: threshold 4 misses input 3.
- Hypothesis: a lower threshold may fix the error.
- Change: threshold 4 → 2.
- Result: mistake count stays at 1 because input 2 becomes wrong.
- Next step: test a more specific candidate or ask whether one threshold is expressive enough.
Debug the right layer
Not every wrong prediction is a model-setting problem. The cause may be:
- input: a feature value is missing or incorrect;
- label: the expected answer was recorded incorrectly;
- feature: important information is absent;
- model: the predictor is too simple;
- evaluation: the metric or test cases do not match the real goal;
- software: the code does not implement the intended rule.
Ask which layer the evidence points toward before changing the model.
Protect final evaluation evidence
If you repeatedly tune a threshold after reading the same held-out results, those examples start guiding model selection. They are no longer a clean final test.
For this tiny teaching exercise we reuse the held-out examples so you can see the effect clearly. In a real workflow, repeated model selection normally uses training/validation evidence, while a final test stays protected until the end.
Evaluation evidence should not secretly become training information.
A reusable debugging record
| Field | What to write |
|---|---|
| Failure | Which example or behavior is wrong? |
| Evidence | What exactly did you observe? |
| Hypothesis | What possible cause are you testing? |
| Change | What one important factor will change? |
| Fixed | What stays the same? |
| Before | Original result |
| After | Result after the change |
| Baseline | Simple reference result |
| Conclusion | What does the evidence support? |
| Next step | What would you test next? |
This pattern will stay useful even when later models become much larger.
Before the Level Project: choose your workspace
The next activity is the Level 0 Project, Data Detective. Local Python is not a hidden Level 0 prerequisite, so you may choose either path: