Level 2: Neural Networks from First Principles
Level 1 showed how to train and evaluate classical machine-learning models fairly. Level 2 opens the model and asks:
What is a neural network actually calculating, and how does an error change the numbers inside it?
You will answer that by working with very small networks whose intermediate values fit on the page.
Start with one limitation
Imagine four points arranged like the corners of a square. Two opposite corners belong to one class and the other two corners belong to the other class.
One straight line cannot separate those two groups correctly.
A neural network can first transform the original inputs into new intermediate values and then make a decision from those transformed values. That extra transformation is useful because some patterns are not easy to express with one simple straight-line rule.
The first lesson uses a tiny pattern called XOR to make this limitation visible. You do not need to know XOR before starting; the four examples are shown explicitly.
A few words you will meet
Each word will be taught with numbers when it becomes useful. This table is only a map.
| Term | Plain meaning in this Level |
|---|---|
| neuron | a small calculation that combines inputs using adjustable weights and a bias |
| layer | a group of neuron calculations that produces the next set of values |
| activation function | a rule applied after a weighted sum; a nonlinear activation lets a network represent patterns that stacked straight-line operations cannot |
| forward pass | calculating from input toward prediction |
| loss | one number that summarizes how wrong the prediction is according to the chosen objective |
| gradient | information about how changing a value would change the loss nearby |
| backpropagation | a method for sending that change information backward through the calculations so parameters can be updated |
A parameter is still the same idea from Level 1: an adjustable model value learned from data. Neural networks simply have many parameters connected through layers.
What you will learn
By the end of this Level, you should be able to:
- explain why a nonlinear network can represent patterns one linear rule cannot;
- compute a neuron's weighted sum and bias by hand;
- compare common activation functions and explain why nonlinearity matters;
- track array and tensor shapes through dense layers;
- perform and debug a two-layer forward pass;
- connect prediction error to a scalar loss;
- interpret derivatives as local sensitivity;
- explain backpropagation as assigning responsibility for error through a chain of calculations;
- carry out a simple parameter update from an explicit gradient;
- reason about mini-batches, optimizers, initialization, and regularization;
- read training curves and diagnose a failure from evidence rather than random tweaking.
No calculus course is assumed. A derivative will first be treated as a simple question:
If this value changes a little, what happens to the result?
Only after that idea is concrete will we use compact derivative notation.
How to learn the math in this Level
For new mathematical ideas, use this order whenever possible:
- Trace a tiny numerical example.
- Predict a direction or shape.
- Run or calculate the operation.
- Explain why the result follows.
- Only then connect it to compact notation or framework code.
A formula is useful when it compresses reasoning you already understand. It should not replace the reasoning.
The learning path
Part 1 — What a network computes
L2.1–L2.5 move from the limits of one linear rule to one neuron, activation functions, layer shapes, and a complete forward pass.
Part 2 — How error reaches parameters
L2.6–L2.9 connect loss, backpropagation intuition, just-in-time derivatives, and one explicit backward/update step.
After L2.9 — Backpropagation in Code, complete Level 2 Mini Checkpoint — Forward and Backward. You should be able to explain the computation without saying only “the framework knows the gradient.”
Part 3 — How training behaves in practice
L2.10–L2.15 cover mini-batches, optimizers, initialization, regularization, training curves, and a structured debugging workshop.
How to use the Labs
Every numbered lesson has a deterministic browser Lab.
When a Lab appears, focus on the named intermediate values. A final neural-network output can look plausible even when an earlier shape, weighted sum, activation, gradient, or update is wrong.
For this Level, a strong explanation usually includes at least one concrete piece of evidence such as:
- a shape;
- a weighted contribution;
- an activation value;
- a gradient sign;
- an update direction;
- a training/validation curve pattern.
What mastery looks like
Do not stop at “I recognize the term.” Try to reach three levels:
- Explain: say the idea in ordinary language.
- Trace: show where it appears in a small computation.
- Transfer: predict what changes in a new tiny case.
Read Quick Check feedback even after a correct answer. The feedback should connect the answer back to the computation.
Level Project
After L2.15 — Neural Network Debugging Workshop, complete Neural Network From Scratch.
The goal is not merely to make a network train. You should be able to explain its shapes, forward computation, loss, gradients, updates, and at least one debugging path from evidence to fix.
Completion is stored locally on this device.