EN
Back to the archive

The encyclopedia · R&D & Science · Technical decision · 1997

LSTM's memory cell let neural networks bridge a thousand time steps

Hochreiter and Schmidhuber's gated cell kept error flowing, letting RNNs learn long sequences others could not.

Technische Universität München / IDSIA

The solution

Recurrent neural networks of the 1980s could in theory learn sequences, but training with backpropagation through time made errors shrink exponentially — the vanishing gradient — so anything needing more than about ten steps was hopeless. Hochreiter's 1991 analysis pinned the problem down.

In 1997, Sepp Hochreiter and Jürgen Schmidhuber published 'Long Short-Term Memory' in Neural Computation. The fix was a memory cell whose self-connection passes error backward at constant strength, with multiplicative input and output gates that learn when to store and release information.

The architecture bridged time lags beyond 1000 steps, even with noisy inputs, at O(1) computational cost per step and weight. In comparisons against RTRL, BPTT and other recurrent learners, LSTM had many more successful runs and solved tasks no previous algorithm could.

Why it worked

  • Constant error flow stops the vanishing gradient
  • Gates decide what to remember, instead of hoping
  • Local computation keeps training practical
  • It solved time lags beyond 1000 steps
What it achievedA self-loop carries error without decayinspired

What can be applied

If information must survive a long journey, build a protected channel with constant gradient and let learned gates control access — do not expect the network to rediscover persistence at every step.

Aftermath

LSTM became the dominant recurrent architecture for speech recognition, translation and handwriting before attention models arrived, and the memory-cell idea remains at the heart of modern sequence networks.

Sources

spotted an error? The archive wants to know.

Related cases