Skip to main content
L1.6

Gradient Descent

Goal

By the end of this lesson, you can describe gradient descent as repeated parameter updates toward lower loss and trace several updates on a one-parameter model.

How do we know which way to change a parameter?​

Suppose our model has one adjustable value, w, and the loss is:

loss = (w - 3)²

The loss is smallest when w = 3.

If we start at w = 0, we want the parameter to move upward. If we start at w = 5, we want it to move downward.

In a simple problem like this, we can see the answer. A real model may have thousands or millions of adjustable parameters, so trying every possible setting is not practical.

Gradient descent gives us a local rule for choosing a direction.

The gradient tells us the local uphill direction​

For the toy loss (w - 3)², the gradient is:

2(w - 3)

You do not need to derive that expression yet. For this Level, treat it as a small direction calculator: plug in the current w, read the sign, and use that sign to decide which way loss rises locally. In Level 2, you will build the derivative idea from small changes and see where formulas like this come from.

At w = 0, the gradient is -6.

Gradient descent uses this update:

new_w = old_w - learning_rate × gradient

With learning rate 0.1:

new_w = 0 - 0.1 × (-6) = 0.6

Subtracting a negative value moves w upward, toward 3.

At a value above 3, the gradient becomes positive, so subtracting it moves w downward.

Follow the same parameter update

The target is w = 3, matching the worked example above. Change the learning rate, then step through the path and watch both w and loss.

Step 0 · x = 0.00 · loss = 9.000
0parameter xloss

A useful picture is standing on a foggy hill. You cannot see the whole landscape, but you can feel the local slope. The gradient points uphill, so gradient descent takes a step in the opposite direction.

Follow several updates in the Lab​

The Lab starts with start_w = -4.0 and repeatedly minimizes the same bowl-shaped loss.

  1. Click Run without editing anything.
  2. Inspect first step:, final w:, and final loss:. The first gradient should be negative, so the first update moves w upward toward 3.
  3. Find the line start_w = -4.0.
  4. Change only -4.0 to 5.0.
  5. Before running, evaluate the sign of 2(w - 3) at w = 5: it is positive. Predict which direction the first update should move.
  6. Click Run and inspect first step:. Confirm that the positive gradient makes gradient descent move w downward toward 3.
  7. Restore start_w = -4.0.

Loading lab…

After the guided comparison, you may try another starting value. Keep target_w, learning_rate, and steps fixed so the starting point is the only changed factor.

The goal is to understand each update, not just to see a final small number.

Direction is only half the problem​

The gradient tells us a local direction. We still need to decide how large a step to take. That is the job of the learning rate, which is the next lesson.

If steps are too large, the parameter can jump across a low-loss region and become unstable. If they are too small, progress can be painfully slow.

Also remember that decreasing training loss only shows that optimization is working on the training objective. It does not prove that the model generalizes to held-out data.

Debug optimization by looking at the path​

When training behaves strangely, inspect the sequence rather than only the final result:

  • Is the loss decreasing, flat, oscillating, or exploding?
  • Are parameter updates moving in the direction you expected?
  • Is the sign in the update rule correct?
  • Is the gradient formula or automatic-gradient computation correct?
  • Is the learning rate appropriate for the scale of the problem?

These questions turn “training failed” into specific evidence you can investigate.

Quick Check

1. What does the gradient describe?
2. Why does gradient descent subtract the gradient?
3. Which evidence best shows whether optimization is moving, stalling, or overshooting?

0 of 3 questions answered.

Key Takeaways

  • Gradient descent repeatedly updates parameters using local loss information.
  • The gradient points toward increasing loss; gradient descent moves in the opposite direction.
  • You can use the gradient formula here without deriving it yet; Level 2 builds that derivative idea from small changes.
  • The learning rate controls how far each update moves.
  • Training-loss progress is optimization evidence, not proof of held-out performance.

Next Lesson

Next, you will focus on the number that controls the size of every gradient step: the learning rate.

References

Lesson actions

Completion is stored locally on this device.

View progress