Splitting Data Without Cheating
Goal
By the end of this lesson, you can create a train/test split, explain why test examples must stay outside training, and use a fixed random seed to make a comparison reproducible.
Practice questions and a final check have different jobs
Suppose you practice with one stack of questions and take a final exam with a different stack.
During practice, it is fine to inspect mistakes and change how you study. The practice questions are helping you learn.
The final exam has a different purpose: it checks whether what you learned works on questions that did not guide your studying.
Machine learning uses the same separation:
- training data is allowed to influence the fitted model;
- test data is held out so it can provide a more independent evaluation later.
If the same examples appear on both sides, the evaluation partly asks whether the model can handle cases it already used while learning.
A tiny split
Imagine 12 numbered examples. A 75/25 split should put 9 examples in training and 3 in testing.
The specific IDs can vary when the split is random, but two rules should hold:
- training and test IDs should not overlap;
- the split should be reproducible when we are comparing models under the same conditions.
A random seed is a starting value used by a pseudo-random process. Using the same seed with the same data and split procedure makes the same random choices again.
The seed does not make the split automatically good. It makes the random part repeatable.
Run the split Lab
The Lab creates 12 examples and calls train_test_split with test_size=0.25 and random_state=42.
- Before running, predict the train and test counts for the 75/25 split: 9 training examples and 3 test examples.
- Click Run.
- Read
train ids:,test ids:, andoverlap:. Confirm thatoverlap:is[]. - Find this argument in the split call:
test_size=0.25
- Change only
0.25to0.50. - Before running, predict that 6 of the 12 examples will now be held out for testing and 6 will remain for training.
- Click Run and count the IDs on each side. Confirm that
overlap:is still empty. - Restore
test_size=0.25before moving on.
Loading lab…
The empty overlap is evidence that no row is literally present in both sets. It is necessary, but it is not the whole fairness story.