본문으로 건너뛰기
L1.16

Boosting: Learning From Previous Errors

Goal

By the end of this lesson, you can explain gradient boosting as a sequence of small learners that correct remaining error, and contrast that process with the diversified-tree aggregation used by Random Forest.

Random Forest and boosting solve the one-tree problem differently​

Random Forest says: train many deliberately different trees, then aggregate their predictions.

Boosting says: start with a simple prediction, inspect what is still wrong, then add another weak learner that improves the remaining error.

The final model is a sum of many small corrections.

Build intuition with residuals​

Suppose a first model predicts 10 when the true value is 14. The residual is 4.

A later learner can learn a correction that pushes the combined prediction upward for examples with similar remaining error.

After that correction, there is a new set of residuals. The next learner focuses on what is still not explained.

For classification the mathematics differs, but the same beginner intuition holds: each stage improves the current ensemble's mistakes.

Weak learners make controlled corrections​

Gradient boosting commonly uses small decision trees as individual learners.

A shallow tree may be weak by itself, but a sequence of small trees can build a flexible model.

The learning rate controls how strongly each new tree's correction is added. A smaller rate usually means each stage contributes less and more stages may be needed.

Watch error shrink stage by stage​

  1. Run the toy gradient-boosting regressor.
  2. Read mean squared error after each stage.
  3. Inspect one example's residual after successive stages.
  4. Increase the number of estimators while keeping depth and learning rate fixed.
  5. Predict whether training error should usually decrease as correction stages are added.
  6. Run again, then change only the learning rate and compare how strongly each stage changes the fit.

Loading lab…

More stages can eventually overfit, so falling training error is not enough to choose the final model. Validation evidence still matters.

The contrast to remember​

Random ForestGradient Boosting
Trees are diversified with row and feature randomnessTrees are added sequentially
Trees do not primarily exist to correct the previous treeEach new learner targets remaining error
Predictions are aggregated across treesPredictions are built as staged corrections
Diversity is centralError correction is central

Both are tree ensembles, but their training logic is different.

Practical descendants​

XGBoost, LightGBM, and CatBoost are important practical gradient-boosting systems with additional engineering and modeling features.

At this level, do not memorize library-specific knobs. First understand the shared idea: sequentially add weak learners that reduce the current model's errors.

Quick Check

1. What does a later boosting learner focus on?
2. What is the central contrast with Random Forest?
3. Why is lower training error after more boosting stages not sufficient evidence?

0 of 3 questions answered.

Key Takeaways

  • Gradient boosting builds a model as a sequence of small corrections.
  • Later learners focus on errors left by the current ensemble.
  • Random Forest relies on diverse trees plus aggregation instead.
  • Learning rate controls the influence of each boosting stage.
  • More stages can reduce training error and still hurt generalization.

Next Lesson

Next, you will leave supervised prediction briefly and learn how K-means and PCA search for structure when no target label is provided.

References

Lesson actions

Completion is stored locally on this device.

View progress