본문으로 건너뛰기
L0.8

Errors: When a Model Gets It Wrong

Goal

By the end of this lesson, you can identify prediction errors, count them, inspect which examples failed, and use an error to form a reasonable debugging question.

An error is more than a bad score​

A prediction error happens when the model's prediction does not match the target label.

On a tiny dataset, we can see every error directly.

Suppose a threshold predictor says:

predict True when the input is at least 3.

And we have these examples:

InputLabel
1False
2False
3True
4False
5True
6True

The rule predicts:

InputPredictionLabelError?
1FalseFalseNo
2FalseFalseNo
3TrueTrueNo
4TrueFalseYes
5TrueTrueNo
6TrueTrueNo

There is one error: input 4.

We could summarize that as 1 mistake out of 6 examples. But the summary is not the whole story.

The identity of the wrong example matters.

One wrong example can suggest several possible causes​

Why did input 4 fail?

From this tiny table alone, several explanations are possible:

  1. The model is too simple. Maybe one threshold cannot describe the real pattern.
  2. A useful feature is missing. Perhaps input 4 needs another piece of information.
  3. The label is noisy or wrong. Maybe the recorded target is mistaken.
  4. The input itself is bad. Perhaps the value was measured incorrectly.

Notice what we should not do:

immediately decide that the label must be wrong because the model disagrees with it.

The model is not the judge of truth. A disagreement is evidence that something deserves investigation.

Counts tell us how much; examples help tell us why​

Imagine two models:

  • Model A makes 3 mistakes.
  • Model B also makes 3 mistakes.

Are they equally good?

Maybe, but not necessarily.

Model A could fail on three harmless edge cases. Model B could fail on the three cases that matter most to users.

Even at Level 0, this gives us an important habit:

Look at the error count, but also inspect the actual errors.

Later you will learn confusion matrices, precision, recall, group-level analysis, and other tools that make error analysis more systematic.

Try the browser Lab​

The Lab below uses the six examples shown above.

  1. Click Run without changing anything.
  2. In stdout, find:
error count: 1
  1. Then find the line beginning with:
wrong example:
  1. Confirm that the wrong example has input 4, prediction True, and label False.

Loading lab…

Now make one controlled model change.

  1. Find:
threshold = 3
  1. Change it to:
threshold = 5
  1. Leave the examples unchanged.
  2. Click Run again.
  3. Compare both the error count and the wrong example with the first run.

Something interesting happens: the total error count is still 1, but the identity of the wrong example changes.

With threshold 5, input 3 becomes the mistake instead of input 4.

This is a useful lesson: the same score can hide different behavior.

Why “fix the label” is not a default debugging move​

You could edit the label for input 4 from False to True. Then the original threshold would look perfect.

But unless you have external evidence that the label was recorded incorrectly, that would be cheating.

Labels are evidence. They can contain errors, but changing them should require a reason such as:

  • checking the original source;
  • finding a data-entry mistake;
  • getting a trusted human review;
  • or discovering that the labeling rule was inconsistent.

A model disagreement is a reason to investigate the label, not proof that the label is wrong.

From error to hypothesis​

A useful debugging note can be very small:

Evidence: input 4 is predicted True but labeled False.

Possible cause: the threshold may be too low.

Test: raise the threshold while keeping the examples fixed.

Observation: the old error disappears, but input 3 becomes wrong.

Interpretation: raising the threshold moves the boundary; it does not solve every case.

This way of thinking turns errors into information.

Different errors can have different costs​

For now, we count every error as one mistake.

In real systems, errors often have different consequences.

For example:

  • a music recommender suggesting one bad song may be annoying but minor;
  • a medical system missing an urgent case can be far more serious;
  • a spam filter blocking an important school email may matter more than allowing one harmless advertisement through.

Later, evaluation will consider not just how many errors happen, but which kinds happen and what they cost.

Quick Check

1. When is a prediction an error?
2. Two settings both make one mistake. What else should you inspect?
3. A model disagrees with one label. What is the best first conclusion?

0 of 3 questions answered.

Key Takeaways

  • An error is a mismatch between prediction and target label.
  • Error counts summarize behavior, but individual errors reveal more detail.
  • The same error count can come from different failed examples.
  • A wrong case can suggest problems with data, labels, features, or the model.
  • Do not change labels just to make a model look correct.
  • Good debugging starts with evidence and a testable hypothesis.

Next Lesson

To test a debugging hypothesis clearly, you need to know what changed between two runs. Next you will practice one of the most useful habits in experimentation: changing one important thing at a time.

References

Lesson actions

Completion is stored locally on this device.

View progress