본문으로 건너뛰기
L0.5

Training Data and Test Data

Goal

By the end of this lesson, you can explain why training data and test data have different jobs, and you can recognize when a test has accidentally started influencing the model choice.

Why we need two different jobs for data​

Suppose you are practicing for a spelling quiz.

During practice, it is fine to look at examples, notice mistakes, and change how you study. The practice questions are helping you improve.

Now imagine the final quiz. If someone shows you every final answer before you take it, your final score no longer tells us much about how well you learned to spell new words.

Machine learning has a similar problem.

We need some examples that are allowed to influence our choices, and some examples that are kept separate so they can give us a more honest check afterward.

That is the basic reason for training data and test data.

  • Training data is used to choose or learn the predictor.
  • Test data is held back and used to evaluate the chosen predictor.

The words sound technical, but the idea is simple: learn here, check there.

A tiny example​

Imagine we are predicting whether a number belongs to the positive class.

Our training examples are:

InputLabel
1False
2False
4True
5True

We decide to use a threshold rule:

predict True when the input is at least the threshold.

From the training examples, threshold 4 looks perfect:

  • 1 -> False
  • 2 -> False
  • 4 -> True
  • 5 -> True

But a good training result is not the end of the story.

We saved two different examples for testing:

InputLabel
3False
6True

Now we can ask whether the threshold chosen from the training data also behaves well on examples that did not choose it.

Why testing on the training data is too easy​

If a predictor is chosen specifically because it fits the training examples, then checking it on those same examples mostly tells us:

“Yes, the choice we made for these examples fits these examples.”

That can be useful, but it does not answer the more important question:

“How well does this choice handle examples that did not guide the choice?”

This is one reason machine learning cares so much about held-out evaluation.

A model that memorizes every training example could have a perfect training score and still perform badly on new cases.

You will study that problem more formally later as overfitting. For now, remember that a training score and a test score answer different questions.

What test leakage means​

A test stops being a clean final check when its results start guiding the model choice.

For example:

  1. Try threshold 3 and look at the test score.
  2. Try threshold 4 and look at the test score.
  3. Try threshold 5 and look at the test score.
  4. Keep whichever threshold had the best test result.

At this point, the test data has influenced the choice. It has quietly become part of the tuning process.

This does not mean that looking at test results is forbidden forever. It means that once test results guide changes, those examples are no longer an untouched final check for the changed system.

Later, you will learn about a third split called validation data, which is often used for tuning while test data stays untouched until the end.

Try the browser Lab​

The Lab below makes the separation visible. You do not need to understand every function yet.

  1. Click Run without changing anything.
  2. Look for these three lines in stdout:
chosen using training data: ...
training mistakes: ...
test mistakes: ...
  1. Notice that the code chooses chosen_threshold using only train_examples.
  2. The starter should choose threshold 4, with 0 training mistakes and 0 test mistakes.

Now make one conceptual change: move one labeled example from training to testing.

  1. In train_examples, remove (4, True).
  2. Add (4, True) to test_examples.
  3. Do not change any other values.
  4. Click Run again.
  5. Compare all three output lines with the first run.

Loading lab…

With (4, True) removed from training, the best threshold chosen from the remaining training examples can change. The test set is also different, so the new test result answers a different question.

That is not a “bad” experiment. It shows why a data split is part of the experiment definition. If two people use different splits, their scores may not be directly comparable.

A subtle point: moving data changes the evidence​

Beginners sometimes think the training/test split is just a bookkeeping detail.

It is not.

Which rows are used for learning can change the model that gets selected. Which rows are held out can change how difficult the final evaluation is.

That is why serious ML experiments record the split or the rule used to create it.

A common misconception​

“If my test score is low, I should keep adjusting the model until the same test score becomes high.”

That is tempting, but every adjustment uses information from the test. Eventually you can overfit your decisions to the test itself.

A cleaner workflow is:

  • use training data to learn;
  • use validation data when you need repeated tuning;
  • save a test set for a final check.

At Level 0, you do not need to build all three sets every time. You only need to understand why evaluation data should not secretly become training information.

Quick Check

1. What is the main job of training data?
2. What is the main job of test data?
3. You repeatedly change a threshold because of its test score. What has happened?

0 of 3 questions answered.

Key Takeaways

  • Training data and test data have different jobs.
  • Training data may influence the predictor; held-out test data checks it.
  • A perfect training score does not guarantee good behavior on new examples.
  • If test results guide repeated tuning, the test is no longer untouched.
  • The exact data split is part of the experiment and should be recorded.

Next Lesson

You now have inputs, labels, and a fairer way to separate learning from checking. Next you will put those pieces together into your first complete AI experiment.

References

Lesson actions

Completion is stored locally on this device.

View progress