Changing One Thing at a Time
Goal
By the end of this lesson, you can design a small comparison where one important factor changes while the rest stay fixed, and explain why that makes the result easier to interpret.
Why controlled changes matter
Suppose you have a predictor that makes 4 mistakes.
You change three things:
- the threshold;
- two labels in the dataset;
- and the test examples.
Now the new run makes only 1 mistake.
Did the new threshold help?
Maybe.
Did correcting the labels help?
Maybe.
Was the new test set simply easier?
Also maybe.
The result improved, but the comparison does not tell us why.
That is the problem controlled experiments try to reduce.
A useful beginner rule is:
When you want to learn the effect of one change, keep the other important parts fixed.
A simple two-run comparison
Imagine these labeled examples:
| Input | Label |
|---|---|
| 1 | False |
| 2 | False |
| 3 | True |
| 4 | True |
| 5 | True |
We compare two threshold settings:
- first threshold:
4 - second threshold:
3
Everything else stays the same:
- same examples;
- same labels;
- same prediction rule;
- same mistake-count metric.
With threshold 4, input 3 is predicted False, so there is 1 mistake.
With threshold 3, all five examples are predicted correctly, so there are 0 mistakes.
Because the threshold is the main difference between the two runs, it is reasonable to say:
On these examples, changing the threshold from 4 to 3 reduced the measured mistake count from 1 to 0.
That sentence is specific. It does not claim that threshold 3 is best everywhere.
“One thing” means one important experimental factor
Sometimes one conceptual change requires editing more than one character or line.
For example, moving one example from training data to test data may require removing it from one list and adding it to another. That is still one experimental change: the data split changed.
The goal is not to count keystrokes. The goal is to know which important factor changed between runs.
Useful factors to track include:
- model settings;
- data version;
- train/test split;
- features;
- metric;
- random seed;
- code version.
At Level 0, we will usually change only one of these at a time.
Try the browser Lab
The Lab below records a comparison explicitly.
- Click Run without changing anything.
- Read the printed dictionary.
- Find these fields:
changedkept_fixedbeforeafterbefore_mistakesafter_mistakes
- Confirm that the record says the threshold changed while the examples and metric stayed fixed.
- The starter should show
before_mistakes: 1andafter_mistakes: 0.
Loading lab…
Now test a different threshold without changing the examples.
- Find:
second_threshold = 3
- Change only
3to2. - Click Run again.
- Compare
after_mistakeswith the first run.
With threshold 2, input 2 is predicted True even though its label is False, so the second setting now makes 1 mistake.
The experiment record makes it easy to say exactly what happened:
- threshold
4: 1 mistake; - threshold
2: 1 mistake; - same examples and metric.
This comparison does not show an improvement, but it still gives useful evidence.
Why keeping records matters
Without an experiment record, it is easy to forget what changed.
Imagine running ten experiments and writing only the scores:
3
2
2
1
4
1
0
2
1
1
Which setting produced the zero? Which dataset was used? Did the metric change? Was one run performed after fixing a bug?
The scores alone are not enough.
Even a small record helps:
change: threshold 4 -> 3
fixed: examples, labels, metric
result: mistakes 1 -> 0
Later, real ML tools will track many more details automatically, but the reasoning habit starts here.
Controlled does not mean perfectly certain
Changing one thing at a time makes a comparison clearer, but it does not magically prove a universal cause.
A tiny dataset may be unrepresentative. Randomness may affect a larger model. Measurement may be noisy. The metric may miss something important.
So controlled comparison gives us better evidence, not absolute certainty.
That distinction is important in science and engineering generally.
A debugging example
Suppose your model suddenly gets worse after a code change.
A useful debugging strategy is:
- Return to the last known working version.
- Apply one suspected change.
- Run the same evaluation.
- Record the result.
- Repeat with the next suspected change if needed.
This is much easier than making five more changes and hoping the problem disappears.
Quick Check
Key Takeaways
- Controlled changes make comparisons easier to explain.
- “One thing” means one important experimental factor, not one keystroke.
- Record both what changed and what stayed fixed.
- A result that does not improve is still useful evidence.
- Controlled experiments reduce ambiguity; they do not create absolute certainty.
Next Lesson
A lower mistake count sounds good, but compared with what? Next you will learn why even a simple predictor should be compared with a meaningful baseline.
References
- Google for Developers, Introduction to Machine Learning.
Completion is stored locally on this device.