Pooling and Downsampling
Goal
By the end of this lesson, you can compute max and average pooling on small patches, explain the information tradeoff of downsampling, and identify when aggressive compression removes task-relevant location detail.
Pooling replaces several spatial values with a summary
Take this 2×2 feature patch:
1 1
1 9
Max pooling returns the largest value:
9
Average pooling returns:
(1 + 1 + 1 + 9) / 4 = 3
Both turn four numbers into one. They simply preserve different summaries.
Max pooling says, “Was there a strong activation somewhere in this patch?”
Average pooling says, “What was the average activation across this patch?”
Neither pooled value tells you the exact original four values afterward. Max pooling preserves the strongest value directly, while average pooling preserves the patch average.
Downsampling trades spatial resolution for compactness
Shrinking a feature map can:
- reduce computation;
- make small shifts less important;
- let later features summarize larger regions.
But the cost is lost detail.
If a task requires exact location—such as locating a tiny defect or segmenting a boundary—aggressive pooling can discard the very information the output needs.
Compare max and average pooling in the Lab
The Lab downsamples a 4×4 feature map to 2×2 outputs.
- Run the starter once and observe that the pooling TODO leaves the pooled values wrong.
- Complete the one-line rule inside
pool2: use the patch maximum formode=="max"and the patch mean otherwise. - Run again and confirm the max/average checks pass.
- Change one patch from all ones to
[1,1;1,9]. - Before running, predict the max-pooled and average-pooled values.
- Run and compare, then reset the starter.
Loading lab…
Now imagine repeated pooling until only one scalar remains. Some global summary may survive, but exact spatial arrangement cannot.
Same pooled output does not mean same original patch
These two patches both max-pool to 9:
9 0 9 9
0 0 9 9
Max pooling cannot distinguish their very different amount or layout of activation.
This example makes the information loss concrete rather than treating pooling as a free compression trick.
Downsampling trades detail for a smaller representation
Take a 2×2 patch:
1 7
3 2
Max pooling returns 7. Average pooling would return 3.25.
Both reduce four values to one, but they preserve different summaries.
Max pooling answers something like “was a strong activation present here?” Average pooling answers “what was the average activation in this region?”
Smaller spatial maps change later computation
Suppose a feature map is 32×32. Downsampling by a factor of 2 gives 16×16.
That reduces the number of spatial positions later layers must process, which can reduce computation and enlarge the effective area of the original image represented by one later feature.
But information was discarded. Two different 2×2 patches can produce the same pooled value.
Pooling is therefore not a free speed trick. It deliberately creates invariance to some local detail.
A useful question is: which differences should the task ignore, and which differences must remain visible?
Quick Check
Key Takeaways
- Pooling is a lossy spatial summary.
- Max and average pooling preserve different information.
- Downsampling can make representations cheaper and less position-sensitive.
- Compression can harm tasks that need precise location or fine detail.
Next Lesson
Next, you will change training images on purpose and decide which transformations preserve the target label.
References
- PyTorch, nn.MaxPool2d.
Completion is stored locally on this device.