Skip to main content
L1.10

Classification

Goal

By the end of this lesson, you can distinguish classification from regression, compute classification accuracy for a simple rule, and inspect which examples are misclassified instead of reporting only one score.

Sometimes the answer is a category, not a quantity​

Regression asks questions such as:

  • How many minutes will the trip take?
  • What will the temperature be?
  • How much energy will the building use?

The output is a number on a continuous scale.

Classification asks questions such as:

  • Is this message spam or not spam?
  • Which species is this flower?
  • Should this transaction be review or no review?

The output is a class or category.

A classifier may still compute a continuous score internally. A decision rule then turns that score into a final class.

A threshold can turn a number into a class​

Suppose a simple school rule says:

predict pass when practice score is at least 60; otherwise predict fail.

The input is numerical, but the output is categorical.

If a score changes from 58 to 62, it crosses the threshold and the predicted class changes.

This separation—score first, decision second—will become important when you study logistic regression and decision thresholds.

Accuracy is easy to compute, but it is not the whole story​

Classification accuracy is:

number correct / number evaluated

If 8 of 10 predictions match their labels, accuracy is 0.8, or 80%.

But imagine a fraud dataset with 99 normal transactions and 1 fraud case. A classifier that always says “normal” gets 99% accuracy while missing every fraud case.

So a useful evaluation asks both:

  • How many predictions were correct overall?
  • Which examples and which classes were wrong?

Run the threshold classifier​

The Lab uses practice scores and a hand-written threshold classifier.

  1. Click Run.
  2. Read the three output lines: predictions: [0, 0, 0, 1, 1, 1, 1, 1], accuracy: 0.75, and mistake indices: [2, 4]. Index 2 is the score 58 (true label 1, predicted 0); index 4 is the score 68 (true label 0, predicted 1).
  3. Find threshold = 60. The scores closest to it are 58 and 62, and the next ones are 68 and 74.
  4. Change only 60 to 70.
  5. Before running, predict: which scores are now below the threshold that were above it before?
  6. Click Run. accuracy is still 0.75, but mistake indices is now [2, 3]. The score 68 stopped being a false alarm, while the score 62 became a new miss.
  7. Press Reset afterward.

The same accuracy hides two different sets of mistakes. That is why you should inspect which examples are wrong, not only how many.

Loading lab…

A changed threshold can improve one set of examples while harming another. Looking only at the final accuracy can hide that tradeoff.

A common misconception​

“Classification means the model only outputs labels.”

Many classifiers first produce a score or probability-like value. The final category appears only after applying a rule such as “positive if score is at least 0.5.”

Keeping the score and the decision separate lets you change decision policy without necessarily refitting the model.

Transfer the idea​

Suppose a smoke detector produces a risk score from 0 to 1 and sounds an alarm above a threshold.

Lowering the threshold may catch more real fires, but it may also create more false alarms. The next lessons will give you language and metrics for reasoning about that tradeoff.

Quick Check

1. What kind of output defines classification?
2. What does accuracy count?
3. Why can high accuracy be misleading?

0 of 3 questions answered.

Key Takeaways

  • Classification predicts categories; regression predicts quantities.
  • A continuous score can be turned into a class by a decision rule.
  • Accuracy is the fraction of matching predictions and labels.
  • Class balance and the identities of mistakes matter alongside accuracy.

Next Lesson

Next, logistic regression will learn a smooth score between 0 and 1 for binary classification.

References

Lesson actions

Completion is stored locally on this device.

View progress