Skip to main content
L1.14

Decision Trees, Bias, Variance, and Overfitting

Goal

By the end of this lesson, you can trace how a decision tree makes recursive splits, compare shallow and deep trees, and use training plus held-out evidence to explain underfitting, overfitting, and tree instability.

Start with a visible decision process​

Imagine deciding whether an outdoor event needs a tent. You might ask whether rain chance is above 40%, then whether wind speed is above 30 km/h.

A decision tree learns questions like these from data. At each node, it chooses a feature and a threshold that separate the current rows into smaller groups. Each child group can be split again. This repeated process is called recursive splitting.

A final leaf stores the prediction for rows that followed that path.

Tree depth controls how many questions can be chained​

A depth-1 tree can make only one split. It may be too simple to capture an important pattern.

A deeper tree can carve feature space into many smaller regions. That extra flexibility can help, but it can also fit accidental details of the training sample.

This connects trees to bias and variance:

  • a very shallow tree can have high bias and underfit;
  • a very deep tree can have high variance and overfit;
  • useful depth depends on data and evaluation evidence.

Read train and held-out scores together​

A shallow tree scoring 61% on training and 59% on held-out data has a small gap, but both scores are weak.

A deep tree scoring 100% on training and 73% held-out has a much larger gap. That is evidence that training-specific detail may not transfer.

Neither pattern is proof by itself. Leakage, tiny evaluation sets, distribution shift, and label noise can also create strange gaps.

Change tree depth in the Lab​

  1. Run the deterministic noisy classification example.
  2. Compare depth 1, depth 3, and an unrestricted tree on training and held-out data.
  3. Increase the middle tree's maximum depth by one.
  4. Predict what should happen to training accuracy.
  5. Run again and inspect both training and held-out accuracy.
  6. Explain whether the extra flexibility helped generalization or mainly improved training fit.

Loading lab…

The question is not “Which tree fits training best?” It is “Which amount of flexibility is supported by evidence that transfers?”

One tree can be unstable​

Decision trees make a sequence of greedy split choices. A small change in the training sample can change an early split, which can change many later branches.

That makes a single tree easy to inspect but sometimes unstable.

In the next lesson, that weakness becomes motivation for an ensemble: train many deliberately different trees and combine their predictions.

Diagnose the experiment before blaming complexity​

Before changing depth, check whether future information leaked, whether the split represents the future task, whether labels are noisy, whether the held-out set is large enough, and whether the same pattern repeats across cross-validation folds.

Model capacity is one possible explanation, not the first explanation for every disappointing score.

Quick Check

1. How does a decision tree build a prediction rule?
2. What pattern is consistent with overfitting?
3. Why can one decision tree be unstable?

0 of 3 questions answered.

Key Takeaways

  • Decision trees learn recursive feature-and-threshold splits.
  • Shallow trees can underfit; very deep trees can overfit.
  • Training and held-out scores must be read together.
  • A small train/test gap is not useful when both scores are poor.
  • Single trees can be unstable because small data changes can alter their structure.

Next Lesson

Next, you will reduce single-tree instability by combining many diverse trees into a Random Forest.

References

Lesson actions

Completion is stored locally on this device.

View progress