Classification
Goal
By the end of this lesson, you can distinguish classification from regression, compute classification accuracy for a simple rule, and inspect which examples are misclassified instead of reporting only one score.
Sometimes the answer is a category, not a quantity
Regression asks questions such as:
- How many minutes will the trip take?
- What will the temperature be?
- How much energy will the building use?
The output is a number on a continuous scale.
Classification asks questions such as:
- Is this message
spamornot spam? - Which species is this flower?
- Should this transaction be
revieworno review?
The output is a class or category.
A classifier may still compute a continuous score internally. A decision rule then turns that score into a final class.
A threshold can turn a number into a class
Suppose a simple school rule says:
predict
passwhen practice score is at least 60; otherwise predictfail.
The input is numerical, but the output is categorical.
If a score changes from 58 to 62, it crosses the threshold and the predicted class changes.
This separation—score first, decision second—will become important when you study logistic regression and decision thresholds.
Accuracy is easy to compute, but it is not the whole story
Classification accuracy is:
number correct / number evaluated
If 8 of 10 predictions match their labels, accuracy is 0.8, or 80%.
But imagine a fraud dataset with 99 normal transactions and 1 fraud case. A classifier that always says “normal” gets 99% accuracy while missing every fraud case.
So a useful evaluation asks both:
- How many predictions were correct overall?
- Which examples and which classes were wrong?
Run the threshold classifier
The Lab uses practice scores and a hand-written threshold classifier.
- Click Run.
- Read the three output lines:
predictions: [0, 0, 0, 1, 1, 1, 1, 1],accuracy: 0.75, andmistake indices: [2, 4]. Index2is the score58(true label1, predicted0); index4is the score68(true label0, predicted1). - Find
threshold = 60. The scores closest to it are58and62, and the next ones are68and74. - Change only
60to70. - Before running, predict: which scores are now below the threshold that were above it before?
- Click Run.
accuracyis still0.75, butmistake indicesis now[2, 3]. The score68stopped being a false alarm, while the score62became a new miss. - Press Reset afterward.
The same accuracy hides two different sets of mistakes. That is why you should inspect which examples are wrong, not only how many.
Loading lab…
A changed threshold can improve one set of examples while harming another. Looking only at the final accuracy can hide that tradeoff.
A common misconception
“Classification means the model only outputs labels.”
Many classifiers first produce a score or probability-like value. The final category appears only after applying a rule such as “positive if score is at least 0.5.”
Keeping the score and the decision separate lets you change decision policy without necessarily refitting the model.
Transfer the idea
Suppose a smoke detector produces a risk score from 0 to 1 and sounds an alarm above a threshold.
Lowering the threshold may catch more real fires, but it may also create more false alarms. The next lessons will give you language and metrics for reasoning about that tradeoff.
Quick Check
Key Takeaways
- Classification predicts categories; regression predicts quantities.
- A continuous score can be turned into a class by a decision rule.
- Accuracy is the fraction of matching predictions and labels.
- Class balance and the identities of mistakes matter alongside accuracy.
Next Lesson
Next, logistic regression will learn a smooth score between 0 and 1 for binary classification.
References
- scikit-learn, Getting Started.
Completion is stored locally on this device.