Why Neural Networks
Goal
By the end of this lesson, you can explain why one straight-line decision rule cannot represent every pattern and how an extra nonlinear transformation can make some patterns easier to separate.
Start with four tiny cases
Imagine two switches, x1 and x2. Each switch can be off (0) or on (1). We want the output to be 1 only when exactly one switch is on.
| x1 | x2 | target | ordinary-language case |
|---|---|---|---|
| 0 | 0 | 0 | neither switch is on |
| 0 | 1 | 1 | only x2 is on |
| 1 | 0 | 1 | only x1 is on |
| 1 | 1 | 0 | both switches are on |
This pattern is called XOR, short for “exclusive OR.” You do not need to memorize the name. The important part is the four-case pattern above.
If you plot x1 horizontally and x2 vertically, the two target-1 cases sit at opposite corners. The two target-0 cases sit at the other two corners.
Now try to draw one straight line that puts both 1 corners on one side and both 0 corners on the other. You cannot.
A straight divider used to separate classes is called a linear decision boundary. So the useful lesson from XOR is:
some classification patterns cannot be separated correctly by one linear decision boundary.
What a neural network changes
A neural network does not have to make the final decision directly from the original two numbers.
It can first calculate a new set of intermediate numbers. A group of calculations that produces those intermediate values is called a hidden layer. “Hidden” only means that these values are inside the model rather than being the original input or final answer.
The next layer can then make its decision from those new values.
One way to picture the process is:
original coordinates
[x1, x2]
↓
hidden layer creates new values
↓
new representation
↓
output decision
A representation is simply the set of values the next part of the network receives. A useful hidden representation can make a difficult raw pattern easier for a later layer.
Why nonlinearity matters
Suppose every layer only stretches, rotates, adds, or combines values in ways that remain linear.
Even if we stack several such linear transformations, the whole stack can still be reduced to one larger linear transformation. We would still have the same straight-boundary limitation.
That is why neural networks usually place a nonlinear activation function between linear transformations. “Nonlinear” means the rule is not restricted to a straight-line relationship.
One common activation is ReLU. In this course, ReLU means:
negative input → 0
positive input → keep the positive value
For example:
ReLU(-2) = 0
ReLU(3) = 3
That simple bend is enough to make stacked layers capable of representations that one linear map cannot create. You will examine activation functions directly in L2.3 — Activation Functions.
Inspect XOR in the Lab
The Lab compares one fixed linear rule with a small ReLU construction and prints the intermediate hidden values explicitly.
- Before running, read the XOR table above and write the expected outputs for
[0,0],[0,1],[1,0], and[1,1]:[0, 1, 1, 0]. - Click Run.
- Compare
linear predictions:withnetwork predictions:. The linear rule should miss part of XOR while the nonlinear construction should match all four targets. - Read
hidden_1:,hidden_2:, andnetwork scores:.hidden_1andhidden_2are intermediate representation values;network scorescombines them before the final class decision. - Find this line:
hidden_2 = relu(s - 1.0)
- Change only
1.0to0.5. - Before running, predict that
hidden_2will become positive for more inputs, so the downstream scores and predictions will change. - Click Run. Compare
hidden_2:,network scores:, andnetwork predictions:with the first run. - Restore
hidden_2 = relu(s - 1.0).
Loading lab…
The point is not that this hand-built network is how you would solve a real product problem. XOR is useful because there are only four cases, so you can see exactly what extra representational flexibility changes—and how changing the representation can also break a working decision rule.
A common misconception
“Neural networks are better because they have more layers.”
More layers only help when the added calculations provide useful flexibility and can be trained from evidence. A deeper model can also overfit, become harder to debug, or simply be unnecessary for a simple task.
The safe conclusion from this experiment is narrow:
this nonlinear network can represent the four-case XOR pattern in a way one straight-line decision rule cannot.
Quick Check
Key Takeaways
- XOR is a four-case pattern where the target is 1 only when exactly one input is 1.
- One straight-line decision boundary cannot separate every classification pattern.
- A hidden layer creates intermediate values inside the network.
- A representation is the set of values passed to the next computation.
- Nonlinear activations let stacked layers do more than one linear map could do.
- Extra model complexity is useful only when the task and evidence justify it.
Next Lesson
Next, in L2.2 — Neurons as Weighted Sums, you will zoom into one neuron and calculate how its inputs, weights, and bias combine into one number.
References
Completion is stored locally on this device.