Training Data and Test Data
Goal
By the end of this lesson, you can explain why training data and test data have different jobs, and you can recognize when a test has accidentally started influencing the model choice.
Why we need two different jobs for data
Suppose you are practicing for a spelling quiz.
During practice, it is fine to look at examples, notice mistakes, and change how you study. The practice questions are helping you improve.
Now imagine the final quiz. If someone shows you every final answer before you take it, your final score no longer tells us much about how well you learned to spell new words.
Machine learning has a similar problem.
We need some examples that are allowed to influence our choices, and some examples that are kept separate so they can give us a more honest check afterward.
That is the basic reason for training data and test data.
- Training data is used to choose or learn the predictor.
- Test data is held back and used to evaluate the chosen predictor.
The words sound technical, but the idea is simple: learn here, check there.
A tiny example
Imagine we are predicting whether a number belongs to the positive class.
Our training examples are:
| Input | Label |
|---|---|
| 1 | False |
| 2 | False |
| 4 | True |
| 5 | True |
We decide to use a threshold rule:
predict
Truewhen the input is at least the threshold.
From the training examples, threshold 4 looks perfect:
- 1 -> False
- 2 -> False
- 4 -> True
- 5 -> True
But a good training result is not the end of the story.
We saved two different examples for testing:
| Input | Label |
|---|---|
| 3 | False |
| 6 | True |
Now we can ask whether the threshold chosen from the training data also behaves well on examples that did not choose it.
Why testing on the training data is too easy
If a predictor is chosen specifically because it fits the training examples, then checking it on those same examples mostly tells us:
“Yes, the choice we made for these examples fits these examples.”
That can be useful, but it does not answer the more important question:
“How well does this choice handle examples that did not guide the choice?”
This is one reason machine learning cares so much about held-out evaluation.
A model that memorizes every training example could have a perfect training score and still perform badly on new cases.
You will study that problem more formally later as overfitting. For now, remember that a training score and a test score answer different questions.
What test leakage means
A test stops being a clean final check when its results start guiding the model choice.
For example:
- Try threshold
3and look at the test score. - Try threshold
4and look at the test score. - Try threshold
5and look at the test score. - Keep whichever threshold had the best test result.
At this point, the test data has influenced the choice. It has quietly become part of the tuning process.
This does not mean that looking at test results is forbidden forever. It means that once test results guide changes, those examples are no longer an untouched final check for the changed system.
Later, you will learn about a third split called validation data, which is often used for tuning while test data stays untouched until the end.
Try the browser Lab
The Lab below makes the separation visible. You do not need to understand every function yet.
- Click Run without changing anything.
- Look for these three lines in
stdout:
chosen using training data: ...
training mistakes: ...
test mistakes: ...
- Notice that the code chooses
chosen_thresholdusing onlytrain_examples. - The starter should choose threshold
4, with0training mistakes and0test mistakes.
Now make one conceptual change: move one labeled example from training to testing.
- In
train_examples, remove(4, True). - Add
(4, True)totest_examples. - Do not change any other values.
- Click Run again.
- Compare all three output lines with the first run.
Loading lab…
With (4, True) removed from training, the best threshold chosen from the remaining training examples can change. The test set is also different, so the new test result answers a different question.
That is not a “bad” experiment. It shows why a data split is part of the experiment definition. If two people use different splits, their scores may not be directly comparable.
A subtle point: moving data changes the evidence
Beginners sometimes think the training/test split is just a bookkeeping detail.
It is not.
Which rows are used for learning can change the model that gets selected. Which rows are held out can change how difficult the final evaluation is.
That is why serious ML experiments record the split or the rule used to create it.
A common misconception
“If my test score is low, I should keep adjusting the model until the same test score becomes high.”
That is tempting, but every adjustment uses information from the test. Eventually you can overfit your decisions to the test itself.
A cleaner workflow is:
- use training data to learn;
- use validation data when you need repeated tuning;
- save a test set for a final check.
At Level 0, you do not need to build all three sets every time. You only need to understand why evaluation data should not secretly become training information.
Quick Check
Key Takeaways
- Training data and test data have different jobs.
- Training data may influence the predictor; held-out test data checks it.
- A perfect training score does not guarantee good behavior on new examples.
- If test results guide repeated tuning, the test is no longer untouched.
- The exact data split is part of the experiment and should be recorded.
Next Lesson
You now have inputs, labels, and a fairer way to separate learning from checking. Next you will put those pieces together into your first complete AI experiment.
References
- scikit-learn, Common pitfalls and recommended practices.
- Google for Developers, Introduction to Machine Learning.
Completion is stored locally on this device.