본문으로 건너뛰기
L1.7

Learning Rate

Goal

By the end of this lesson, you can explain the learning rate as optimization step size, recognize too-small and too-large rates from training behavior, and design a fair learning-rate comparison.

The same direction can succeed or fail because of step size​

In the previous lesson, the gradient told us which local direction should reduce loss.

But knowing the direction is not enough.

Imagine walking downhill:

  • with tiny steps, you may move safely but very slowly;
  • with useful-sized steps, you can make steady progress;
  • with giant jumps, you may cross the valley, bounce back, or move farther away.

The learning rate controls this step size.

The basic gradient-descent update is:

parameter = parameter - learning_rate × gradient

The gradient decides the direction. The learning rate scales the distance.

Compare two updates before thinking about a whole training run​

Suppose the gradient is -4.

With learning rate 0.01:

change = -0.01 × (-4) = +0.04

With learning rate 0.5:

change = -0.5 × (-4) = +2

Both updates move in the same direction, but the second travels fifty times farther.

That simple calculation explains why changing the learning rate can completely change a training curve even when the gradient direction is unchanged.

Three common training patterns​

A too-small learning rate often looks like:

  • loss decreases;
  • progress is steady;
  • but many steps produce only tiny improvement.

A useful rate often looks like:

  • loss decreases clearly;
  • updates remain stable;
  • progress reaches a low-loss region in a reasonable number of steps.

A too-large rate can look like:

  • loss oscillates;
  • loss suddenly grows;
  • parameter values jump across the useful region;
  • training becomes numerically unstable.

These are clues, not universal proofs. Other bugs can create similar symptoms, so isolate the learning rate when you test it.

Run the learning-rate comparison​

The Lab uses the same one-parameter bowl loss with learning rates 0.02, 0.2, and 1.1.

  1. Click Run.
  2. Read final losses: {0.02: 21.658119, 0.2: 0.001792, 1.1: 1878.542396}. Every run starts at start_w = -4.0 and takes the same steps = 10.
  3. Label each rate: 0.02 is slow (the loss is still large), 0.2 is useful (the loss is almost zero), and 1.1 is unstable (the loss became much larger than where it started).
  4. Find learning_rates = [0.02, 0.2, 1.1]. Change only 1.1 to 0.9. Keep start_w and steps unchanged so the learning rate is the only difference.
  5. Before running, predict: 0.9 is still large, but smaller than 1.1. Will it explode, or settle somewhere?
  6. Click Run. The new rate should finish at about 0.565. It no longer explodes, because each step overshoots the minimum by a little less than it did before, so the zig-zag slowly shrinks. It is still worse than 0.2.
  7. Press Reset afterward.

Loading lab…

Keeping the other conditions fixed matters because otherwise you cannot tell whether the changed behavior came from the learning rate or from something else.

Why one good learning rate is not universal​

A value that works for one problem may fail for another.

Feature scale, model architecture, optimizer, batch size, and loss shape can all change how large an update is sensible.

Modern systems may also use a learning-rate schedule, changing the rate during training. Even then, the core idea remains the same: step size is a training setting that affects optimization and must be recorded if you want someone else to reproduce the run.

A common misconception​

“If training is slow, just use the largest learning rate that does not crash immediately.”

A rate can avoid an obvious crash and still produce noisy or unreliable optimization. Look at the path of the loss, not just whether the program finishes.

Quick Check

1. What does the learning rate control?
2. What can a too-small learning rate cause?
3. Why change one setting at a time in this Lab?

0 of 3 questions answered.

Key Takeaways

  • Learning rate is optimization step size.
  • Too-small rates can be slow; too-large rates can overshoot or become unstable.
  • Training curves reveal learning-rate problems better than one final number.
  • Learning rate is part of the experiment record and should be changed in a controlled comparison.

Next Lesson

Next, you will see why the numerical scale of input features can change optimization behavior even when the underlying information is the same.

References

Lesson actions

Completion is stored locally on this device.

View progress