Skip to main content
L0.10

Fair Comparisons and Baselines

Goal

By the end of this lesson, you can explain what a baseline is, compare a predictor with a simple baseline on the same examples, and tell why a model score needs context before it can be called good.

Is 80% accuracy good?​

Imagine someone says:

“My model is 80% accurate.”

That sounds impressive, but we are missing important information.

Suppose 80% of the examples belong to the same class. A strategy that always guesses that most common class might also get 80% accuracy without learning any useful pattern.

In that case, the model's 80% score is not an improvement over a very simple alternative.

This is why machine learning uses baselines.

A baseline is a simple reference strategy that gives us something meaningful to compare against.

A baseline answers “better than what?”​

A model score by itself does not tell us whether the model is useful.

A baseline gives the score context.

Examples of simple baselines include:

  • always predict the most common class;
  • always predict the average number;
  • use a simple hand-written rule;
  • use the system that is already deployed today.

A baseline does not need to be clever. In fact, its value often comes from being simple and easy to understand.

If a complicated model cannot beat a simple baseline, that is important evidence.

A tiny majority baseline​

Suppose our training examples are:

InputLabel
1False
2False
3False
4True
5True

There are three False labels and two True labels.

The majority label is therefore False.

A majority baseline says:

ignore the input and always predict False.

That sounds almost too simple to be useful, but that is exactly why it is a good reference.

Now evaluate both a threshold model and the baseline on the same held-out examples:

InputTrue label
2False
3False
4True
6True

The threshold model uses threshold 4:

predict True when input >= 4

It gets all four held-out examples correct.

The majority baseline predicts False every time. It gets the first two correct but misses the last two.

So on this held-out set:

  • threshold model: 0 mistakes;
  • majority baseline: 2 mistakes.

Now the model's score has context. It did better than a simple alternative under the same evaluation.

Fair comparison means same test and same metric​

Imagine a model is tested on easy examples and the baseline is tested on harder examples.

Even if the model gets a better score, we do not know whether the model is better or whether its test was easier.

For a fair comparison, keep the evaluation conditions aligned:

  • same held-out examples;
  • same labels;
  • same metric;
  • same rules for counting success.

The thing that should differ is the strategy being compared.

Try the browser Lab​

The Lab below implements exactly this comparison.

  1. Click Run without changing anything.
  2. In stdout, find:
model mistakes: ...
baseline label: ... baseline mistakes: ...
  1. The starter should show:
    • model mistakes: 0
    • baseline label: False
    • baseline mistakes: 2
  2. Notice that both strategies were evaluated on the same test_examples.

Loading lab…

Now make the model slightly worse while leaving the baseline and test set unchanged.

  1. Find:
threshold = 4
  1. Change only 4 to 5.
  2. Click Run again.
  3. Compare model mistakes with baseline mistakes.

With threshold 5, the model misses input 4, so it now makes 1 mistake. It is worse than before, but it still beats the majority baseline's 2 mistakes on this held-out set.

This is a more informative statement than simply saying “the model got 75%.”

What if the model does not beat the baseline?​

That does not automatically mean the project is worthless.

It tells us that the current complexity has not earned its place yet.

Possible reasons include:

  • the feature does not contain enough useful information;
  • the model setting is poor;
  • the dataset is too small or noisy;
  • the baseline is already very strong;
  • the metric does not match what we care about.

A baseline can save time by revealing when a complicated approach is not actually improving the result.

A baseline should represent a real alternative​

Not every easy comparison is useful.

Suppose you create an intentionally terrible baseline that makes random guesses badly. Beating it does not prove much.

A good baseline should answer a realistic question such as:

“Is this model better than doing the simplest sensible thing?”

In a real product, the most important baseline may be the current shipped system. A new model that is exciting in a notebook may not be worth deploying if it is slower, more expensive, and no better than what users already have.

Complexity is not the same as progress​

Machine learning often involves impressive-looking models, large datasets, GPUs, and many parameters.

But complexity by itself is not a result.

A simpler system that is equally accurate, cheaper, faster, and easier to understand may be the better engineering choice.

Baselines help keep us focused on measured improvement instead of novelty.

Quick Check

1. What is a baseline?
2. What makes a model-versus-baseline comparison fair?
3. A complex model cannot beat a simple baseline. What does that tell you?

0 of 3 questions answered.

Key Takeaways

  • A model score needs context.
  • A baseline is a simple reference strategy.
  • Compare the model and baseline on the same evaluation data with the same metric.
  • Beating a weak or unfair baseline does not prove much.
  • Complexity is useful only when it provides a meaningful benefit.
  • Failing to beat a baseline is valuable evidence, not something to hide.

Next Lesson

Once you have a fair comparison, you still need to communicate what happened without making claims that are larger than the evidence. Next you will turn an experiment result into a clear explanation.

References

Lesson actions

Completion is stored locally on this device.

View progress