Skip to main content
L2.7

Backpropagation Intuition

Goal

By the end of this lesson, you can explain backpropagation as local sensitivity passed backward through a dependency chain and use a small perturbation to infer a useful parameter direction.

Start with the question backpropagation must answer​

A network produces a loss at the end of a forward pass.

Training needs to know:

If I change this earlier weight a tiny amount, how will the final loss change?

A large network has many parameters, so trying every combination would be hopelessly expensive.

Backpropagation answers the question efficiently by reusing the computation dependencies from the forward pass.

A small perturbation makes the idea visible​

Suppose one weight is currently w.

Compute the loss at:

  • w - h
  • w + h

for a small h.

If the loss is lower at w + h, then increasing the weight locally appears useful.

If the loss is lower at w - h, decreasing the weight locally appears useful.

This is a finite-difference way to observe gradient direction. It is slow for a real network, but excellent for building intuition and checking gradient code.

Local effects multiply through a chain​

Imagine this dependency:

w -> prediction -> loss

The loss changes when prediction changes, and prediction changes when w changes.

The chain rule combines those two local sensitivities:

dLoss/dw = dLoss/dPrediction × dPrediction/dw

For a longer network, the chain contains more operations, but the reasoning is the same.

If a value branches into several downstream paths, the gradient contributions from those paths add back together.

Backpropagation: forward to the loss, backward to the weight

The forward pass creates a chain from the weight to the loss. Backpropagation follows that chain backward to find how the weight affected the loss.

Stage 1/4 · Forward pass
Forward passweightw = 1× 2prediction2target 5loss 9Backward passgradient = -12update weightnew w = 2.2

Forward pass: prediction = 2 × 1 = 2

The visual uses the derivative for this tiny squared-error example. The next lesson derives the derivative rules behind that calculation; here, focus on the direction of information flow and the sign of the gradient.

Probe one weight in the Lab​

The Lab uses the simple function:

prediction = weight * 2.0
loss = (prediction - 5.0) ** 2

The best weight for this tiny problem is 2.5 because 2.5 × 2 = 5.

  1. Click Run with w = 1.0 and epsilon = 1e-4.
  2. Read finite-difference gradient: and gradient descent should move weight:. The gradient should be negative, so the useful local direction is to increase the weight toward 2.5.
  3. Change only w = 1.0 to w = 3.0.
  4. Before running, reason from the target: 3.0 × 2 = 6, which is now above the target 5. Predict that the gradient sign should become positive and gradient descent should move the weight downward.
  5. Click Run and confirm the direction changes to decrease.
  6. Restore w = 1.0.

Loading lab…

A useful special case about epsilon​

You can also change epsilon = 1e-4 to epsilon = 0.5 and rerun with w = 1.0. In this particular Lab, the centered finite-difference estimate stays exactly the same because the loss is a quadratic function of w. That is a special property of this toy—not a rule that large perturbations are always safe.

For more complicated curved functions, a large epsilon can compare points that are too far apart to represent the slope right around the current value. The word local still matters.

A finite-difference check is evidence about the gradient. It is not the efficient training algorithm itself.

A common misconception​

“Backpropagation sends the final prediction error unchanged to every parameter.”

Different parameters influence the output through different paths and local operations. A parameter multiplied by a zero activation can receive a very different gradient from a parameter on an active path.

That is why local sensitivity matters.

If a gradient sign surprises you, draw the dependency chain.

For each link ask:

  • if the earlier value increases, does the later value increase or decrease?
  • is that local effect positive, negative, or zero?
  • how do the signs combine along the path?

This reasoning is often enough to find a reversed sign before reading a long formula.

Quick Check

1. What does backpropagation combine?
2. What can finite differences check?
3. At w=3 in this Lab, prediction is 6 while the target is 5. Which local direction should reduce loss?

0 of 3 questions answered.

Key Takeaways

  • Backpropagation asks how each earlier value affects the final loss.
  • The chain rule combines local sensitivities through dependency paths.
  • Finite differences make gradient direction visible and can check implementation.
  • A centered difference happens to be exact for this quadratic toy even with a larger epsilon; that does not remove the need for local checks on more complex functions.
  • Gradients are path-specific; the final error is not copied unchanged to every parameter.

Next Lesson

Next, you will learn exactly enough derivative math to describe those local sensitivities directly.

References

Lesson actions

Completion is stored locally on this device.

View progress