Your First AI Experiment
Goal
By the end of this lesson, you can describe and run a small AI experiment with a clear question, a fair comparison, one controlled change, and a result you can explain.
Start with one question
An experiment begins when there is something you want to find out.
Imagine that you have two versions of the same simple predictor. One uses threshold 5. The other uses threshold 4.
You want to know:
Which threshold works better on the examples we set aside for testing?
You could run both versions and compare their scores. But for that comparison to mean anything, the two runs should be treated the same way. They should use the same test examples and the same way of measuring success. The threshold should be the important thing that changes.
Then the result gives you evidence for answering the question.
For example:
Question: Which threshold works better here?
Keep the same:
- test examples
- scoring method
Change:
- threshold 5 → threshold 4
Observe:
- the score for each threshold
Use the result to answer the question.
That is the basic shape of the experiment you will run in this lesson.
Writing model code is only one part of machine learning. We also need to be able to tell what a result actually shows.
If a score goes up, useful questions include:
- What did we change?
- What stayed the same?
- Which examples did we test on?
- How did we measure success?
Without that information, a number such as “accuracy = 0.9” does not tell us very much by itself.
At this level, a small experiment can be as simple as:
- Ask one clear question.
- Keep the important comparison conditions the same.
- Change one setting.
- Measure the result in the same way.
- Explain what the result does—and does not—show.
Our tiny problem: is the fruit ripe?
Imagine that we describe a fruit with one number called a sweetness score.
| Fruit | Sweetness score | Ripe? |
|---|---|---|
| A | 2 | No |
| B | 3 | No |
| C | 5 | Yes |
| D | 6 | Yes |
| E | 7 | Yes |
| F | 4 | No |
We use a very simple rule:
predict ripe when
sweetness >= threshold
The threshold is the part we are allowed to change.
The first four examples will act as training examples. The last two will be held out for our comparison.
Training examples:
| Sweetness | Ripe? |
|---|---|
| 2 | No |
| 3 | No |
| 5 | Yes |
| 6 | Yes |
Held-out examples:
| Sweetness | Ripe? |
|---|---|
| 7 | Yes |
| 4 | No |
Decide what success means before looking at the answer
For this experiment, we will use accuracy:
accuracy = number correct / number tested
Because there are only two held-out examples, each example changes the score by half.
That is a useful warning: this tiny experiment is good for learning the workflow, but two examples are far too few for a strong real-world claim.
Compare two thresholds fairly
We will compare:
- threshold
5 - threshold
4
Before calculating, reason through the two held-out examples.
For input 7, both thresholds predict True, which matches the label.
For input 4:
- threshold
5predictsFalse; - threshold
4predictsTrue.
The true label for input 4 is False.
So you can already predict that threshold 5 should do better on this tiny held-out set.
Notice what makes this comparison useful: the test examples and the metric stay the same. The threshold is the important thing that changes.
Run the experiment in the browser
The Lab below contains exactly this experiment.
- Click Run without changing anything.
- In
stdout, find the two lines beginning withthreshold. - You should see threshold
5score1.0and threshold4score0.5on the same held-out examples. - Translate those decimals back into examples:
1.0means 2 of 2 correct;0.5means 1 of 2 correct.
Loading lab…
The result supports this narrow statement:
On these two held-out examples, using threshold
5gave higher accuracy than using threshold4.
It does not prove that threshold 5 is best for every fruit, every dataset, or every future input.
Make one more controlled change
If you want to extend the experiment:
- Find
threshold_b = 4. - Change only
4to6. - Leave
threshold_a,train,test, and the accuracy function unchanged. - Before pressing Run, predict the new score for threshold
6. - Run and compare the two lines again.
Threshold 6 also gets both held-out examples correct, so this tiny test cannot tell us whether 5 or 6 is better overall.
That is an important kind of result. An experiment does not always produce one clear winner. Sometimes the evidence says, “These two choices look the same on the data we tested.”
Why changing several things at once is confusing
Imagine a different experiment where you:
- change the threshold;
- replace one test example;
- change the accuracy metric;
- and edit a label.
Then the score improves.
Which change caused the improvement?
From that one comparison, you cannot tell.
This is why beginner experiments often follow the rule change one important thing at a time. It does not make the world perfectly controlled, but it makes the evidence much easier to interpret.
Experiments can fail and still be useful
Suppose you predict that a change will help, but the score gets worse.
That is not a failed learning experience. You learned that your hypothesis was not supported by this test.
Good experiment notes keep both successful and unsuccessful runs. Hiding the bad runs makes it harder to understand what really happened.
A simple experiment record
For a small ML experiment, you should be able to write something like this:
| Part | Example |
|---|---|
| Question | Does threshold 5 beat threshold 4? |
| Data held fixed | Same two held-out examples |
| Changed | Threshold only |
| Metric | Accuracy |
| Result | 1.0 vs 0.5 |
| Conclusion | Threshold 5 did better on this held-out set |
| Limitation | Only two held-out examples |
That record is more valuable than a bare number because someone else can understand what the number means.
Quick Check
Key Takeaways
- An ML experiment starts with a clear question and uses measured evidence to answer it.
- Decide what will be measured before interpreting the result.
- Keep important conditions fixed when comparing one change.
- A result supports a specific claim about the data and setup that produced it.
- Tiny datasets are useful for learning but weak evidence for broad claims.
- Unsuccessful hypotheses are useful when they are recorded honestly.
Next Lesson
You have now run a complete tiny experiment. Next we will zoom in on one prediction at a time and separate three things that beginners often mix together: the input, the model's prediction, and the true label.
References
- scikit-learn, Model selection and evaluation.
- scikit-learn, Common pitfalls and recommended practices.
Completion is stored locally on this device.