Skip to main content
L0.9

Changing One Thing at a Time

Goal

By the end of this lesson, you can design a small comparison where one important factor changes while the rest stay fixed, and explain why that makes the result easier to interpret.

Why controlled changes matter​

Suppose you have a predictor that makes 4 mistakes.

You change three things:

  • the threshold;
  • two labels in the dataset;
  • and the test examples.

Now the new run makes only 1 mistake.

Did the new threshold help?

Maybe.

Did correcting the labels help?

Maybe.

Was the new test set simply easier?

Also maybe.

The result improved, but the comparison does not tell us why.

That is the problem controlled experiments try to reduce.

A useful beginner rule is:

When you want to learn the effect of one change, keep the other important parts fixed.

A simple two-run comparison​

Imagine these labeled examples:

InputLabel
1False
2False
3True
4True
5True

We compare two threshold settings:

  • first threshold: 4
  • second threshold: 3

Everything else stays the same:

  • same examples;
  • same labels;
  • same prediction rule;
  • same mistake-count metric.

With threshold 4, input 3 is predicted False, so there is 1 mistake.

With threshold 3, all five examples are predicted correctly, so there are 0 mistakes.

Because the threshold is the main difference between the two runs, it is reasonable to say:

On these examples, changing the threshold from 4 to 3 reduced the measured mistake count from 1 to 0.

That sentence is specific. It does not claim that threshold 3 is best everywhere.

“One thing” means one important experimental factor​

Sometimes one conceptual change requires editing more than one character or line.

For example, moving one example from training data to test data may require removing it from one list and adding it to another. That is still one experimental change: the data split changed.

The goal is not to count keystrokes. The goal is to know which important factor changed between runs.

Useful factors to track include:

  • model settings;
  • data version;
  • train/test split;
  • features;
  • metric;
  • random seed;
  • code version.

At Level 0, we will usually change only one of these at a time.

Try the browser Lab​

The Lab below records a comparison explicitly.

  1. Click Run without changing anything.
  2. Read the printed dictionary.
  3. Find these fields:
    • changed
    • kept_fixed
    • before
    • after
    • before_mistakes
    • after_mistakes
  4. Confirm that the record says the threshold changed while the examples and metric stayed fixed.
  5. The starter should show before_mistakes: 1 and after_mistakes: 0.

Loading lab…

Now test a different threshold without changing the examples.

  1. Find:
second_threshold = 3
  1. Change only 3 to 2.
  2. Click Run again.
  3. Compare after_mistakes with the first run.

With threshold 2, input 2 is predicted True even though its label is False, so the second setting now makes 1 mistake.

The experiment record makes it easy to say exactly what happened:

  • threshold 4: 1 mistake;
  • threshold 2: 1 mistake;
  • same examples and metric.

This comparison does not show an improvement, but it still gives useful evidence.

Why keeping records matters​

Without an experiment record, it is easy to forget what changed.

Imagine running ten experiments and writing only the scores:

3
2
2
1
4
1
0
2
1
1

Which setting produced the zero? Which dataset was used? Did the metric change? Was one run performed after fixing a bug?

The scores alone are not enough.

Even a small record helps:

change: threshold 4 -> 3
fixed: examples, labels, metric
result: mistakes 1 -> 0

Later, real ML tools will track many more details automatically, but the reasoning habit starts here.

Controlled does not mean perfectly certain​

Changing one thing at a time makes a comparison clearer, but it does not magically prove a universal cause.

A tiny dataset may be unrepresentative. Randomness may affect a larger model. Measurement may be noisy. The metric may miss something important.

So controlled comparison gives us better evidence, not absolute certainty.

That distinction is important in science and engineering generally.

A debugging example​

Suppose your model suddenly gets worse after a code change.

A useful debugging strategy is:

  1. Return to the last known working version.
  2. Apply one suspected change.
  3. Run the same evaluation.
  4. Record the result.
  5. Repeat with the next suspected change if needed.

This is much easier than making five more changes and hoping the problem disappears.

Quick Check

1. Why keep most important conditions fixed between two runs?
2. You changed both the threshold and the test examples, then the score improved. What can you safely say?
3. Why record what stayed fixed?

0 of 3 questions answered.

Key Takeaways

  • Controlled changes make comparisons easier to explain.
  • “One thing” means one important experimental factor, not one keystroke.
  • Record both what changed and what stayed fixed.
  • A result that does not improve is still useful evidence.
  • Controlled experiments reduce ambiguity; they do not create absolute certainty.

Next Lesson

A lower mistake count sounds good, but compared with what? Next you will learn why even a simple predictor should be compared with a meaningful baseline.

References

Lesson actions

Completion is stored locally on this device.

View progress