Gradient Descent and Learning Rate
Goal
By the end of this lesson, you can trace gradient-descent updates, explain what the learning rate controls, and diagnose slow, stable, or unstable optimization from a short loss history.
Direction and step size solve different parts of the problem
Suppose a model has one adjustable value, w, and the loss is smallest at w = 3.
If w starts below 3, we want it to move upward. If it starts above 3, we want it to move downward. A real model has many adjustable values, so we need a repeatable rule instead of trying every possible setting.
The gradient tells us the local uphill direction of the loss. Gradient descent moves in the opposite direction:
parameter = parameter - learning_rate × gradient
Two ideas are packed into that line:
- the gradient supplies a local direction;
- the learning rate controls how far the update moves.
Keeping those roles separate makes optimization easier to reason about.
Trace one update
For the toy loss (w - 3)², the gradient is 2(w - 3). You do not need to derive that formula yet.
At w = 0, the gradient is -6. With learning rate 0.1:
new w = 0 - 0.1 × (-6) = 0.6
Subtracting a negative value moves w upward, toward the low-loss point. At w = 5, the gradient is positive, so subtracting it moves w downward.
Direction and step size
The target is w = 3. Change the learning rate and step through the path to see how the same local direction can produce slow, useful, or unstable movement.
Why learning rate matters
Imagine walking downhill in fog. Tiny steps can be safe but slow. Medium steps can make steady progress. Giant jumps can cross the valley and bounce from side to side.
The same pattern appears in optimization:
- a too-small learning rate often decreases loss slowly;
- a useful rate makes clear, stable progress;
- a too-large rate can oscillate or make loss grow.
There is no single learning rate that works for every problem. Feature scale, model shape, optimizer, and loss geometry all affect useful step sizes.
Compare paths in the Lab
The Lab runs the same one-parameter problem with several learning rates.
- Run the starter unchanged.
- Read the first gradient and the final loss for each rate.
- Confirm that the moderate rate reaches a low-loss region faster than the tiny rate.
- Confirm that the largest example rate becomes unstable.
- Change only the largest rate to a slightly smaller value.
- Predict whether the path will still diverge, oscillate while shrinking, or settle smoothly.
- Run again and compare the full loss histories, not only the final number.
Loading lab…
A fair learning-rate experiment keeps the starting point, number of steps, loss, and data fixed. Otherwise a changed result may have more than one explanation.
Optimization evidence is not generalization evidence
A decreasing training loss means the optimizer is improving the chosen training objective. It does not prove that the model will work well on held-out data.
Later lessons will keep these questions separate:
- Did optimization fit the training objective?
- Does the fitted model generalize to new examples?
Both matter.
Debug the path, not just the endpoint
When optimization behaves strangely, inspect the sequence:
- Is loss decreasing, flat, oscillating, or exploding?
- Are updates moving in the expected direction?
- Did only the learning rate change?
- Are feature scales making one direction dominate?
- Does a smaller step stabilize the run?
This turns “training failed” into a testable diagnosis.
Quick Check
Key Takeaways
- Gradient descent repeatedly moves parameters using local loss information.
- The gradient supplies direction; the learning rate supplies step size.
- Too-small rates can be slow, while too-large rates can overshoot or diverge.
- Loss histories reveal optimization behavior better than one final number.
- Better training loss is not automatically better held-out performance.
Next Lesson
Next, you will see why feature scale can change the geometry of optimization and how to fit scaling without leaking evaluation information.
References
- scikit-learn, Stochastic Gradient Descent.
Completion is stored locally on this device.