The encyclopedia · R&D & Science · Technical decision · 1997
LSTM's memory cell let neural networks bridge a thousand time steps
Hochreiter and Schmidhuber's gated cell kept error flowing, letting RNNs learn long sequences others could not.
Technische Universität München / IDSIA
The solution
Recurrent neural networks of the 1980s could in theory learn sequences, but training with backpropagation through time made errors shrink exponentially — the vanishing gradient — so anything needing more than about ten steps was hopeless. Hochreiter's 1991 analysis pinned the problem down.
In 1997, Sepp Hochreiter and Jürgen Schmidhuber published 'Long Short-Term Memory' in Neural Computation. The fix was a memory cell whose self-connection passes error backward at constant strength, with multiplicative input and output gates that learn when to store and release information.
The architecture bridged time lags beyond 1000 steps, even with noisy inputs, at O(1) computational cost per step and weight. In comparisons against RTRL, BPTT and other recurrent learners, LSTM had many more successful runs and solved tasks no previous algorithm could.
Why it worked
- Constant error flow stops the vanishing gradient
- Gates decide what to remember, instead of hoping
- Local computation keeps training practical
- It solved time lags beyond 1000 steps
What can be applied
If information must survive a long journey, build a protected channel with constant gradient and let learned gates control access — do not expect the network to rediscover persistence at every step.
Aftermath
LSTM became the dominant recurrent architecture for speech recognition, translation and handwriting before attention models arrived, and the memory-cell idea remains at the heart of modern sequence networks.
Sources
spotted an error? The archive wants to know.