Errors: When a Model Gets It Wrong
Goal
By the end of this lesson, you can identify prediction errors, count them, inspect which examples failed, and use an error to form a reasonable debugging question.
An error is more than a bad score
A prediction error happens when the model's prediction does not match the target label.
On a tiny dataset, we can see every error directly.
Suppose a threshold predictor says:
predict
Truewhen the input is at least3.
And we have these examples:
| Input | Label |
|---|---|
| 1 | False |
| 2 | False |
| 3 | True |
| 4 | False |
| 5 | True |
| 6 | True |
The rule predicts:
| Input | Prediction | Label | Error? |
|---|---|---|---|
| 1 | False | False | No |
| 2 | False | False | No |
| 3 | True | True | No |
| 4 | True | False | Yes |
| 5 | True | True | No |
| 6 | True | True | No |
There is one error: input 4.
We could summarize that as 1 mistake out of 6 examples. But the summary is not the whole story.
The identity of the wrong example matters.
One wrong example can suggest several possible causes
Why did input 4 fail?
From this tiny table alone, several explanations are possible:
- The model is too simple. Maybe one threshold cannot describe the real pattern.
- A useful feature is missing. Perhaps input
4needs another piece of information. - The label is noisy or wrong. Maybe the recorded target is mistaken.
- The input itself is bad. Perhaps the value was measured incorrectly.
Notice what we should not do:
immediately decide that the label must be wrong because the model disagrees with it.
The model is not the judge of truth. A disagreement is evidence that something deserves investigation.
Counts tell us how much; examples help tell us why
Imagine two models:
- Model A makes 3 mistakes.
- Model B also makes 3 mistakes.
Are they equally good?
Maybe, but not necessarily.
Model A could fail on three harmless edge cases. Model B could fail on the three cases that matter most to users.
Even at Level 0, this gives us an important habit:
Look at the error count, but also inspect the actual errors.
Later you will learn confusion matrices, precision, recall, group-level analysis, and other tools that make error analysis more systematic.
Try the browser Lab
The Lab below uses the six examples shown above.
- Click Run without changing anything.
- In
stdout, find:
error count: 1
- Then find the line beginning with:
wrong example:
- Confirm that the wrong example has input
4, predictionTrue, and labelFalse.
Loading lab…
Now make one controlled model change.
- Find:
threshold = 3
- Change it to:
threshold = 5
- Leave the examples unchanged.
- Click Run again.
- Compare both the error count and the wrong example with the first run.
Something interesting happens: the total error count is still 1, but the identity of the wrong example changes.
With threshold 5, input 3 becomes the mistake instead of input 4.
This is a useful lesson: the same score can hide different behavior.
Why “fix the label” is not a default debugging move
You could edit the label for input 4 from False to True. Then the original threshold would look perfect.
But unless you have external evidence that the label was recorded incorrectly, that would be cheating.
Labels are evidence. They can contain errors, but changing them should require a reason such as:
- checking the original source;
- finding a data-entry mistake;
- getting a trusted human review;
- or discovering that the labeling rule was inconsistent.
A model disagreement is a reason to investigate the label, not proof that the label is wrong.
From error to hypothesis
A useful debugging note can be very small:
Evidence: input 4 is predicted True but labeled False.
Possible cause: the threshold may be too low.
Test: raise the threshold while keeping the examples fixed.
Observation: the old error disappears, but input 3 becomes wrong.
Interpretation: raising the threshold moves the boundary; it does not solve every case.
This way of thinking turns errors into information.
Different errors can have different costs
For now, we count every error as one mistake.
In real systems, errors often have different consequences.
For example:
- a music recommender suggesting one bad song may be annoying but minor;
- a medical system missing an urgent case can be far more serious;
- a spam filter blocking an important school email may matter more than allowing one harmless advertisement through.
Later, evaluation will consider not just how many errors happen, but which kinds happen and what they cost.
Quick Check
Key Takeaways
- An error is a mismatch between prediction and target label.
- Error counts summarize behavior, but individual errors reveal more detail.
- The same error count can come from different failed examples.
- A wrong case can suggest problems with data, labels, features, or the model.
- Do not change labels just to make a model look correct.
- Good debugging starts with evidence and a testable hypothesis.
Next Lesson
To test a debugging hypothesis clearly, you need to know what changed between two runs. Next you will practice one of the most useful habits in experimentation: changing one important thing at a time.
References
- Google for Developers, Introduction to Machine Learning.
Completion is stored locally on this device.