Skip to main content
L2.5

Forward Pass by Hand

Goal

By the end of this lesson, you can perform and verify a two-layer forward pass from inputs to prediction while keeping every important intermediate value visible.

A forward pass follows the arrows​

A network does not learn during a forward pass. It simply computes an output using its current parameters.

For a two-layer network, the path can be written as:

  1. z1 = X @ W1 + b1
  2. a1 = activation(z1)
  3. z2 = a1 @ W2 + b2
  4. output from z2 using the task's final transformation

The names are useful because they separate two different things:

  • z1: the weighted sums before activation;
  • a1: the transformed values passed to the next layer.

If you confuse them, downstream calculations can look reasonable while using the wrong quantity.

Trace one ReLU example​

Suppose a hidden layer produces:

z1 = [-2, 3]

After ReLU:

a1 = [0, 3]

The first hidden unit contributes zero to the next layer because its pre-activation was negative. The second unit passes its positive value onward.

If you change one first-layer bias so -2 becomes +1, then that hidden unit wakes up:

a1 changes from [0, 3] to [1, 3]

and the output layer now receives a new contribution.

This lets you predict which downstream values should change before running code.

Run the forward pass Lab​

The Lab uses one two-feature example, a two-neuron hidden layer, ReLU, and one output neuron.

  1. Click Run and read the outputs in order: z1: [[0.1, -1.45]], a1: [[0.1, 0.0]], and prediction: 0.17. (prediction is the output layer's z2.)
  2. Verify the second hidden pre-activation by hand. With x = [2.0, -1.0], the second column of W1 is [-0.25, 0.75] and its bias is -0.2: 2.0 × (-0.25) + (-1.0) × 0.75 + (-0.2) = -1.45. ReLU turns -1.45 into 0.0.
  3. Find b1 = np.array([0.1, -0.2]). Change only the second bias, -0.2, to 1.5, so that hidden pre-activation becomes positive.
  4. Before running, predict which values must change and which stay fixed. Does the first hidden neuron change?
  5. Click Run. Now z1: [[0.1, 0.25]] and a1: [[0.1, 0.25]]: the second neuron is “on” and passes 0.25 forward. Because its output weight in W2 is -0.7, the prediction falls from 0.17 to -0.005. The first neuron did not change at all.
  6. Press Reset afterward.

Loading lab…

When debugging, stop at the first intermediate value that disagrees with your hand calculation. Every later mismatch may simply be a consequence of that earlier error.

A semantic failure can survive shape checks​

Imagine applying W2 directly to the original input instead of a1.

In a specially chosen toy example, shapes might still happen to fit. The code would run, but the network would have skipped the hidden computation you intended.

Shape correctness is necessary. It is not sufficient. Also verify that each operation receives the right meaningful value.

Forward computation creates the dependency path for learning​

Conceptually, each result depends on earlier values:

X -> z1 -> a1 -> z2 -> loss

Backpropagation will later traverse these dependencies in reverse to ask how a small change in an earlier parameter would affect the final loss.

A forward pass is a chain of named transformations​

Consider a tiny network:

input x = [2, 1]
linear layer → z
ReLU → h
final linear layer → prediction

Do not jump directly from input to prediction. Record each intermediate value.

If the first layer produces z = [-1, 3], ReLU gives:

h = [0, 3]

The final layer therefore receives [0,3], not the original [2,1] and not the pre-activation [-1,3].

That distinction matters when debugging.

Stop at the first disagreement​

Suppose code returns the wrong final prediction.

If your hand-calculated first-layer output already differs from the program, inspecting the final layer is wasted effort. The earliest mismatch is closer to the cause.

This gives a reusable debugging method:

  1. verify the input;
  2. verify each linear pre-activation;
  3. verify the activation output;
  4. continue layer by layer;
  5. only then inspect the final prediction.

Later Transformer debugging uses the same habit with much larger tensors.

Quick Check

1. Does a forward pass update parameters?
2. Where should you stop when debugging a forward mismatch?
3. What does the forward pass create conceptually for backpropagation?

0 of 3 questions answered.

Key Takeaways

  • A forward pass computes predictions from current parameters; it does not update them.
  • Keep pre-activations and activations distinct.
  • Intermediate values make the computation auditable.
  • Debug from the earliest mismatch.
  • The forward dependency chain is the path that backpropagation will reverse.

Next Lesson

Next, you will turn prediction quality into a scalar loss that gives training a measurable objective.

References

Lesson actions

Completion is stored locally on this device.

View progress