Skip to main content
L3.4

Pooling and Downsampling

Goal

By the end of this lesson, you can compute max and average pooling on small patches, explain the information tradeoff of downsampling, and identify when aggressive compression removes task-relevant location detail.

Pooling replaces several spatial values with a summary​

Take this 2×2 feature patch:

1 1
1 9

Max pooling returns the largest value:

9

Average pooling returns:

(1 + 1 + 1 + 9) / 4 = 3

Both turn four numbers into one. They simply preserve different summaries.

Max pooling says, “Was there a strong activation somewhere in this patch?”

Average pooling says, “What was the average activation across this patch?”

Neither pooled value tells you the exact original four values afterward. Max pooling preserves the strongest value directly, while average pooling preserves the patch average.

Downsampling trades spatial resolution for compactness​

Shrinking a feature map can:

  • reduce computation;
  • make small shifts less important;
  • let later features summarize larger regions.

But the cost is lost detail.

If a task requires exact location—such as locating a tiny defect or segmenting a boundary—aggressive pooling can discard the very information the output needs.

Compare max and average pooling in the Lab​

The Lab downsamples a 4×4 feature map to 2×2 outputs.

  1. Run the starter once and observe that the pooling TODO leaves the pooled values wrong.
  2. Complete the one-line rule inside pool2: use the patch maximum for mode=="max" and the patch mean otherwise.
  3. Run again and confirm the max/average checks pass.
  4. Change one patch from all ones to [1,1;1,9].
  5. Before running, predict the max-pooled and average-pooled values.
  6. Run and compare, then reset the starter.

Loading lab…

Now imagine repeated pooling until only one scalar remains. Some global summary may survive, but exact spatial arrangement cannot.

Same pooled output does not mean same original patch​

These two patches both max-pool to 9:

9 0 9 9
0 0 9 9

Max pooling cannot distinguish their very different amount or layout of activation.

This example makes the information loss concrete rather than treating pooling as a free compression trick.

Downsampling trades detail for a smaller representation​

Take a 2×2 patch:

1 7
3 2

Max pooling returns 7. Average pooling would return 3.25.

Both reduce four values to one, but they preserve different summaries.

Max pooling answers something like “was a strong activation present here?” Average pooling answers “what was the average activation in this region?”

Smaller spatial maps change later computation​

Suppose a feature map is 32×32. Downsampling by a factor of 2 gives 16×16.

That reduces the number of spatial positions later layers must process, which can reduce computation and enlarge the effective area of the original image represented by one later feature.

But information was discarded. Two different 2×2 patches can produce the same pooled value.

Pooling is therefore not a free speed trick. It deliberately creates invariance to some local detail.

A useful question is: which differences should the task ignore, and which differences must remain visible?

Quick Check

1. What does pooling necessarily do?
2. Which pooling keeps one extreme activation most directly?
3. What is a risk of repeated aggressive pooling?

0 of 3 questions answered.

Key Takeaways

  • Pooling is a lossy spatial summary.
  • Max and average pooling preserve different information.
  • Downsampling can make representations cheaper and less position-sensitive.
  • Compression can harm tasks that need precise location or fine detail.

Next Lesson

Next, you will change training images on purpose and decide which transformations preserve the target label.

References

Lesson actions

Completion is stored locally on this device.

View progress