본문으로 건너뛰기
L2.10

Batching and Mini-Batches

Goal

By the end of this lesson, you can distinguish full-batch and mini-batch gradients, explain why different mini-batches produce different gradient estimates, and recognize sum-versus-mean scaling mistakes.

One example gives one view of the gradient​

Suppose four training examples each produce a gradient contribution for the same parameter:

[-2, -1, 1, 4]

The full-batch mean gradient is:

(-2 - 1 + 1 + 4) / 4 = 0.5

That uses every example before taking an update.

Now split the examples into two mini-batches:

  • batch A: [-2, -1] → mean -1.5
  • batch B: [1, 4] → mean 2.5

The two mini-batches point in different directions even though their average across the full dataset is positive.

That is not a bug. Each batch sees only part of the training evidence.

Why use mini-batches?​

Real datasets are often too large to process all at once.

Mini-batches:

  • use less memory per update;
  • allow efficient vectorized computation;
  • produce more frequent parameter updates;
  • introduce variation in gradient estimates because each subset is different.

That variation is often called gradient noise. It can affect optimization behavior and interacts with learning rate and batch size.

Sum versus mean changes update scale​

Suppose a batch has four gradients all equal to 2.

Their sum is 8; their mean is 2.

If one implementation sums and another averages but both use the same learning rate, the update magnitude changes with batch size.

So when batch-size changes unexpectedly alter training, check whether the loss and gradients are normalized consistently.

Compare mini-batch gradients in the Lab​

The Lab computes per-example gradients and compares the full-data mean with two mini-batch means.

  1. Click Run and read per-example gradients: [-2.0, -8.0, -18.0, -32.0]. Each number says how one example alone would push the weight.
  2. Check the full-batch mean by hand: (-2 - 8 - 18 - 32) / 4 = -15. The output shows full-batch gradient: -15.0.
  3. Read mini-batch gradients: [-5.0, -25.0]. The first mini-batch averages the first two examples; the second averages the last two. Their average, (-5 - 25) / 2, is again -15.
  4. Now use four mini-batches of one example each. Replace the whole mini_batch_gradients = [...] block with:
mini_batch_gradients = [float(g) for g in per_example_gradients]
  1. Before running, predict: will the individual batch gradients be closer together or more spread out than [-5.0, -25.0]?
  2. Click Run. The batch gradients become [-2.0, -8.0, -18.0, -32.0], a range of 30 instead of 20. Smaller batches give noisier individual update directions, even though their average still matches the full batch.
  3. Press Reset afterward.

Loading lab…

Batch dimension and feature dimension have different jobs​

In many tensor conventions, the leading dimension indexes examples:

(batch, features)

Changing batch size should not silently change the number of features or learned weight dimensions.

This is the same shape discipline from L2.4 — Layers and Shapes, now applied to training groups.

A common misconception​

“Mini-batch gradients are wrong because they differ from the full-batch gradient.”

They are estimates based on fewer examples. They can still be useful update directions, and repeated batches collectively expose the optimizer to the full dataset.

Quick Check

1. Why can mini-batch gradients differ?
2. Why does sum-vs-mean matter?
3. Which dimension usually indexes examples in this level?

0 of 3 questions answered.

Key Takeaways

  • Full-batch gradients use all training examples before an update.
  • Mini-batches estimate the gradient from subsets and therefore vary.
  • Sum-versus-mean conventions change update magnitude.
  • Batch size is separate from feature or parameter shape.
  • Batch size interacts with optimizer behavior and learning rate.

Next Lesson

Next, you will compare how different optimizers turn the same kind of gradient evidence into parameter updates over time.

References

Lesson actions

Completion is stored locally on this device.

View progress