Level 3 Mini Checkpoint: Represent, See, and Remember
Complete this checkpoint after L3.7 — Recurrent Models and Memory. The goal is to check whether you can reason about information flow, not merely recognize CNN/RNN vocabulary.
For every answer, say what information is kept, changed, or lost.
Part 1 — Representation or raw input?
A model turns a 5×5 image into two values: vertical-edge score and horizontal-edge score.
- Why are the two values a representation?
- Name one kind of information they keep.
- Name one kind of information they may discard.
- Give a second task for which the discarded information could matter.
Part 2 — Convolution by hand
For the patch
1 0
1 0
and filter
1 -1
1 -1
compute the output value.
Then move the same local edge one column to the right in a larger image. Explain what weight sharing predicts about the feature-map response.
Part 3 — Pooling tradeoff
For [1,1;1,9]:
- compute max pooling;
- compute average pooling;
- explain which one preserves the existence of one strong activation more directly;
- name information neither pooled value can reconstruct exactly.
Part 4 — Augmentation semantics
For an up vs down arrow classifier, decide whether these can keep the original label:
- small brightness change;
- horizontal flip;
- vertical flip;
- 180-degree rotation.
Defend each answer from label meaning, not from a memorized list of “standard augmentations.”
Part 5 — Sequence order
Compare [1,2,5] with [5,2,1].
- What statistic is identical?
- What simple feature has opposite sign?
- Give a real sequence task where sorting values would destroy useful information.
Part 6 — Recurrent memory trace
Using h_t = tanh(0.7*h_(t-1) + x_t) and h_0=0, predict the sign of the state after [1,0,0].
Then explain why the first input can still affect the third state despite the later inputs being zero.
Transfer check
Imagine the same recurrence but with recurrent weight 0. Which parts of your explanation change? Which concept stays the same?
A strong answer says that a hidden state still exists computationally, but previous state no longer contributes to the new one in this toy recurrence.
Check your reasoning after you try
- The two edge scores are a representation because they summarize the image for a purpose. They keep edge evidence but discard many exact pixel details.
- The hand convolution is
1×1 + 0×(-1) + 1×1 + 0×(-1) = 2. Weight sharing means the same local pattern can trigger a similar response at another location. - Max pooling of
[1,1;1,9]is 9; average pooling is 3. Max keeps the strongest activation more directly, while neither result lets you reconstruct all four original values. - Brightness change and horizontal flip can preserve an up/down label in many datasets. Vertical flip and 180° rotation swap up with down, so they are unsafe unless the task defines them differently.
- Order-sensitive features distinguish
[1,2,5]from[5,2,1]even though summaries such as the sum match. - With recurrent weight
0.7, the first positive input can keep affecting later hidden states. With recurrent weight0, that path is cut after the current step.
Pass condition
You are ready for L3.8 — Why Long Context Is Hard when you can explain:
- what a representation keeps and loses;
- one convolution output from actual numbers;
- one pooling information tradeoff;
- one safe/unsafe augmentation from target semantics;
- why order matters;
- how recurrent state carries earlier influence.
If you cannot yet explain one item in ordinary language, revisit the matching browser Lab and make one controlled change before continuing.
Completion is stored locally on this device.