Skip to main content
L3.9

LSTM and Gated Memory

Goal

By the end of this lesson, you can predict how changing an LSTM gate changes memory and explain how LSTM/GRU improve recurrent memory while still leaving a sequential path that motivates Attention.

One running note is easy to overwrite​

Imagine reading:

Mina put the red key in the top drawer. Several unrelated events happen. Much later, someone asks where the red key is.

A simple RNN keeps updating one hidden state at every step. New input and old state are mixed into a new state, then the old state is gone.

L3.08 made one consequence visible with numbers: if only 80% of an earlier influence survives each step, then after ten steps 0.8¹⁰ ≈ 0.11. Repeated transformations can make distant information and gradients fade; repeated factors above one can make them grow too much.

That can work well for nearby information. But if the model must protect one early fact while many later inputs arrive, repeatedly rewriting the same state creates a difficult job. Gated recurrence is one response to that measured problem.

A gated recurrent model asks more specific questions.

old memory
|
forget gate -> what should be removed?
input gate -> what should be written?
|
new memory
|
output gate -> what should be exposed now?

Before looking at equations, predict the behavior:

  • If an old fact is still useful, the forget decision should keep most of it.
  • If the current input is irrelevant, the input/write decision should add little.
  • If stored information is not needed for the current output, the output/read decision can hide much of it without necessarily erasing the memory.

Those are the core LSTM ideas.

Memory and hidden state do different jobs​

A standard LSTM carries two related values through time:

  • the cell state is the longer-lived memory path;
  • the hidden state is the exposed representation used by the current step and passed onward.

This separation is useful intuition, not a claim that cell state is permanent storage. Both are finite learned vectors, and both can lose information.

The important design change is that the model no longer has to completely rewrite one exposed state at every step.

A simplified gated update​

Once the gate roles are clear, a simplified scalar memory update is:

new_memory = forget_gate * old_memory
+ input_gate * candidate

exposed_state = output_gate * tanh(new_memory)

The actual LSTM computes learned vector-valued gates from the current input and previous hidden state. The gates usually pass through a sigmoid, so values near 0 block a path and values near 1 let more through.

For the forget gate, a value near 1 means “keep most of the old memory.” A value near 0 means “erase most of it.”

For the input gate, a value near 1 lets more candidate information be written.

For the output gate, a value near 1 exposes more of the current memory through the hidden state.

You do not need to memorize the full LSTM equations yet. You should be able to predict what happens when one of these controls changes.

Change a gate and watch the consequence​

The Lab compares a simple recurrent trace with a simplified gated-memory trace. An early signal appears first, followed by several distractor steps.

  1. Run the Lab unchanged. Compare simple_trace with gated_trace.
  2. Find forget_gate = 0.95. Predict what will happen to the early signal if you change it to 0.20.
  3. Run again. The gated memory should lose the early signal much faster.
  4. Restore forget_gate = 0.95, then change input_gate = 1.0 to 0.0. Predict whether the first signal can be written into memory.
  5. Restore the input gate, then set output_gate = 0.0. Notice that the stored memory can remain non-zero while the exposed hidden value becomes zero.
  6. Restore all defaults when you finish.

Loading lab…

This Lab deliberately keeps each gate at one fixed scalar so you can isolate what that control does. A real LSTM recomputes vector-valued gate values at every time step from the current input and previous hidden state.

The point is not that a high forget gate is always correct. The useful gate values depend on the sequence and task. The point is that the model can learn separate controls for preserving, writing, and exposing information.

Why the memory path can help learning​

In a simple recurrent chain, useful information and its training signal repeatedly pass through the same kind of nonlinear state rewrite.

An LSTM creates a more direct additive memory path. When the forget gate stays near 1, information and gradients can travel through many steps with less repeated shrinking than in a basic recurrence.

That helps with long dependencies, but it does not solve them perfectly. Gates can still learn the wrong behavior, memory is still finite, and the computation still advances step by step.

A GRU, or gated recurrent unit, uses the same broad idea—learn when to preserve old information and when to replace it—but with a simpler state design.

A GRU commonly uses:

  • an update gate to balance old state against new candidate information;
  • a reset gate to control how much previous state contributes while forming the candidate.

Unlike the usual LSTM description, a GRU does not maintain a separate cell-state path plus hidden state in the same way.

For this Level, the important comparison is:

ModelMain idea
simple RNNrepeatedly rewrite one recurrent hidden state
LSTMadd separate gated keep/write/read decisions and a cell-state memory path
GRUuse a simpler gated recurrent state with fewer gate/state distinctions

Both LSTM and GRU are still recurrent: step t depends on state produced by earlier steps.

The bridge to Attention​

The sequence-model story now has a clear progression:

simple RNN
-> distant-information problem
-> LSTM / GRU
-> improved gated memory
-> still sequential
-> Attention
-> Transformer

Gating improves the path through recurrence. Attention changes the path itself: a later position can directly combine information from selected earlier positions instead of requiring every useful signal to survive only through one recurrent chain.

That is why LSTM/GRU are not a failed detour. They solve an important problem in recurrence and make the motivation for Attention easier to see.

Quick Check

1. What does an LSTM forget-gate value near 1 mean in the simplified memory path?
2. Why distinguish cell state from hidden state in LSTM intuition?
3. What important limitation remains for LSTMs and GRUs?
4. How does a GRU differ from the LSTM picture used in this lesson?

0 of 4 questions answered.

Key Takeaways

  • A simple RNN repeatedly overwrites one recurrent state.
  • LSTM separates longer-lived cell-state memory from the currently exposed hidden state.
  • Forget, input, and output gates control keep, write, and read/expose decisions.
  • Gated additive paths can improve information and gradient flow without creating perfect memory.
  • GRU is a related, simpler gated recurrent architecture.
  • Gated recurrence is still sequential, which sets up the motivation for Attention.

Next Lesson

Next, you will turn discrete items into learned vectors and ask what geometric similarity means inside a representation.

References

Lesson actions

Completion is stored locally on this device.

View progress