Level 2 Mini Checkpoint: Forward and Backward
Complete this checkpoint after L2.9 — Backpropagation in Code. This checkpoint does not get a new L2.x number because it reviews the forward/backward ideas before the Level moves on to batching and optimizers.
For every part, try to give both the answer and the reasoning path. If you can only remember a formula but cannot explain what its numbers mean, revisit the matching Lab.
Part 1 — Shapes before numbers
A batch contains 5 examples with 3 input features each. A layer has 4 neurons.
- What is the input shape?
- What shape must the weight matrix have for
X @ W + b? - What is the output shape?
A strong answer names the roles of the dimensions: examples, input features, and output neurons.
Part 2 — Predict a forward pass
For a neuron with inputs [2, -1], weights [0.5, 2.0], and bias 1.0, compute the weighted sum.
Then apply ReLU.
Explain whether ReLU changes the value and point to the sign of the pre-activation as your reason.
Part 3 — Loss is evidence
Two models predict a binary target of 1:
- Model A predicts
0.9. - Model B predicts
0.55.
Which should have the lower binary cross-entropy loss? Explain using “probability assigned to the true outcome,” not just a memorized metric name.
Part 4 — Derivative as local sensitivity
For f(w) = (w - 3)^2, explain what a negative derivative at w = 1 means.
Do not start from a rule. Describe which tiny direction should reduce the function and why.
Part 5 — Backpropagation story
Write 4–6 sentences that explain backpropagation without saying “the computer just knows the gradient.” Include:
- a final loss;
- a chain of earlier computations;
- local derivatives;
- the chain-rule idea;
- how a parameter receives a useful direction for change.
Transfer check
Now imagine the same ideas in a network with three inputs instead of two. Which parts of your reasoning change, and which stay the same?
You should recognize that the arrays become larger, but the same pattern remains: forward dependencies, scalar loss, local sensitivities, chained gradients, parameter update.
Check your reasoning after you try
- Shapes:
Xis(5, 3),Wis(3, 4), and the layer output is(5, 4). The first dimension counts examples; the others count input and output features. - Forward pass:
2×0.5 + (-1)×2.0 + 1.0 = 0. ReLU leaves the value at 0. - Loss: Model A should have lower binary cross-entropy because it assigns more probability to the true outcome
1. - Derivative sign: at
w=1, moving a little to the right moves toward the minimum at 3, so the local slope is negative. - Backpropagation: the final loss sends local sensitivity backward through the recorded computation chain. Each parameter gets a gradient that tells how a small change would affect that loss.
Pass condition
You are ready for L2.10 — Batching and Mini-Batches when you can:
- predict the shapes before running code;
- calculate one forward path;
- explain why one prediction has lower loss;
- reason about derivative sign;
- tell the backpropagation story in ordinary language.
If one part is unclear, revisit L2.4 — Layers and Shapes through L2.9 — Backpropagation in Code and rerun the relevant browser Lab with one controlled change.
Completion is stored locally on this device.