Metrics Beyond Accuracy
Goal
By the end of this lesson, you can read a confusion matrix, compute precision, recall, and F1, and choose which metric deserves attention based on the real cost of different mistakes.
Two models can have the same accuracy and fail differently
Imagine a medical screening system.
There are two kinds of wrong positive-class decisions:
- false positive: the model raises an alarm for a person who is actually negative;
- false negative: the model misses a person who is actually positive.
Both count as one mistake in accuracy. Their consequences may be very different.
That is why classification evaluation often starts with the confusion matrix, which keeps the mistake types separate.
Start with counts before formulas
For the positive class, suppose we have:
TP = 8true positives;FP = 2false positives;FN = 4false negatives.
Precision asks:
When the model predicted positive, how often was it right?
precision = TP / (TP + FP) = 8 / 10 = 0.8
Recall asks:
Of all actual positives, how many did the model find?
recall = TP / (TP + FN) = 8 / 12 ≈ 0.67
If missing a real positive is especially costly, recall deserves close attention because every false negative lowers it.
If false alarms are costly because a human team can investigate only a few cases, precision may matter strongly because false positives lower it.
F1 combines precision and recall, but it does not choose the goal for you
F1 is the harmonic mean of precision and recall. It becomes high only when both are reasonably high.
That is useful when you want one summary balancing the two, but F1 does not know the real cost of a false positive versus a false negative.
A single metric never removes the need to understand the decision.
Trace one changed prediction in the Lab
The Lab uses this fixed prediction array:
y_pred = np.array([1, 1, 0, 1, 1, 0, 0, 0, 0, 0])
The third prediction is 0 while the third true label is 1, so that example is a false negative.
- Click Run and read the confusion matrix first, then accuracy, precision, recall, and F1.
- In
y_pred, change only the third value from0to1:
y_pred = np.array([1, 1, 1, 1, 1, 0, 0, 0, 0, 0])
- Before running, predict the count change: one false negative should become one true positive. No false-positive count should change.
- Click Run. Check the confusion matrix before the summary metrics and confirm the FN→TP move.
- Then inspect precision, recall, accuracy, and F1 and explain why each changed in the direction it did.
- Restore the third prediction to
0.
Loading lab…
After this guided pass, choose one different prediction to flip. State which confusion-matrix cell should change before you run it.
Working from the counts makes the metrics easier to understand and much easier to debug.
Thresholds create metric tradeoffs
For many classifiers, lowering the positive decision threshold creates more positive predictions.
That may:
- turn some false negatives into true positives, increasing recall;
- also turn some true negatives into false positives, decreasing precision.
The model is not necessarily “better” or “worse” in one universal sense. The operating point changed.
A responsible choice connects the threshold and metrics to the consequences of the decision.
Look beyond the overall confusion matrix when needed
An overall metric can still hide important variation.
Depending on the task, you may also need:
- metrics by subgroup;
- threshold curves;
- probability calibration;
- cost-weighted evaluation;
- performance over time or under distribution shift.
The central habit is the same as Level 0: a summary is useful, but inspect the evidence hidden underneath it.
Quick Check
Key Takeaways
- Accuracy hides which kinds of classification mistakes occurred.
- Precision focuses on the correctness of positive predictions.
- Recall focuses on finding actual positives.
- F1 balances precision and recall but does not encode every real-world cost.
- Choose metrics and thresholds from the decision consequences, not from habit.
Next Lesson
Next, you will repeat training and validation across several folds so one lucky split does not dominate model selection.
References
- scikit-learn, sklearn.metrics.
Completion is stored locally on this device.