본문으로 건너뛰기
L2.4

Layers and Shapes

Goal

By the end of this lesson, you can predict the input, weight, bias, and output shapes of a dense layer and diagnose a shape mismatch from dimension meaning rather than trial-and-error transposes.

Shapes tell you what an array means​

Suppose a batch contains 3 examples and each example has 2 input features.

Then the input matrix has shape:

X: (3, 2)

Now suppose the next layer has 4 neurons.

Using the convention in this level:

  • X: (batch, input_features) = (3, 2)
  • W: (input_features, output_features) = (2, 4)
  • b: (output_features,) = (4,)
  • X @ W + b: (batch, output_features) = (3, 4)

The result has one row per example and one column per output neuron.

Follow the dense-layer interfaces

Select each stage and read what every dimension means before checking whether the multiplication is legal.

Input X: (3, 2)
B = 3
three examples in the batch
F_in = 2
two input features per example

Why the inner dimensions must match​

In X @ W, each row of X contains 2 feature values. Each output neuron's weights must therefore contain 2 matching input weights.

That is why the inner dimensions are both 2:

(3, 2) @ (2, 4) -> (3, 4)

If you transpose W to (4, 2), you would try:

(3, 2) @ (4, 2)

The inner dimensions 2 and 4 disagree, so the multiplication does not make sense.

Changing batch size should not change learned parameter shapes​

If the same model receives 5 examples instead of 3:

  • X becomes (5, 2);
  • W remains (2, 4);
  • b remains (4,);
  • output becomes (5, 4).

Batch size counts examples. It is not a learned feature dimension.

Verify the shapes in the Lab​

The Lab sends three two-feature examples through a four-neuron dense layer.

  1. Before running, write the expected shapes of X, W, b, and the output Z.
  2. Click Run and compare with the two output lines: X (3, 2) W (2, 4) b (4,) and Z (3, 4) A (3, 4).
  3. Now make the batch bigger. Find X = np.array([[1.0, 2.0], [0.0, -1.0], [3.0, 0.5]]) and add two more two-feature examples:
X = np.array([[1.0, 2.0], [0.0, -1.0], [3.0, 0.5], [2.0, 2.0], [-1.0, 0.0]])
  1. Before running, predict which shapes should change (only the batch size) and which should stay fixed (the layer's W and b).
  2. Click Run. You should see X (5, 2) and Z (5, 4) A (5, 4), while W (2, 4) and b (4,) are unchanged. The layer's parameters do not depend on how many examples you send at once.
  3. Press Reset afterward.

Loading lab…

When a framework later reports a mismatch such as (32, 128) @ (64, 10), translate the numbers into meanings: “the data currently has 128 features but the layer expects 64.”

A dangerous debugging habit​

“Transpose or flatten until the error disappears.”

That can make the code run while mixing examples, features, or channels incorrectly.

Instead:

  1. write a semantic name beside each dimension;
  2. identify which dimensions the operation requires to match;
  3. fix the upstream representation or parameter shape that violates that meaning.

Trace dimensions before tracing numbers​

Suppose a batch contains 5 examples, each with 3 input features.

X shape = (5, 3)

A linear layer with 4 output units needs a weight matrix compatible with those feature dimensions. Conceptually:

3 input features → 4 output features

After the layer:

output shape = (5, 4)

The batch count stays 5. The feature width changes from 3 to 4.

If the next layer expects 4 inputs and produces 2 outputs, the shape becomes (5,2).

This simple shape trace often catches bugs before any arithmetic is inspected.

Shape-valid does not always mean meaning-valid​

A tensor can have the expected dimensions but the wrong axis meaning.

For example, (5,3) might mean “5 examples × 3 features,” while another operation accidentally treats it as “5 time steps × 3 channels.”

The dimensions still fit some calculations, but the semantics are wrong.

When writing shape notes, label axes with names such as B for batch and C for features instead of recording only raw numbers.

Quick Check

1. What must match in `X @ W`?
2. If batch size changes, which learned layer parameter shape must change?
3. What is a good shape-debugging habit?

0 of 3 questions answered.

Key Takeaways

  • Shape is part of a tensor's meaning.
  • Dense-layer multiplication requires matching input-feature dimensions.
  • Batch size can change without changing learned parameter shapes.
  • Debug shape errors by naming dimensions before reshaping or transposing.

Next Lesson

Next, you will use those shapes to trace a complete two-layer forward pass from input to prediction.

References

Lesson actions

Completion is stored locally on this device.

View progress