Learning to forget: Continual prediction with LSTM

被引:2413
作者
Gers, FA [1 ]
Schmidhuber, J [1 ]
Cummins, F [1 ]
机构
[1] IDSIA, CH-6900 Lugano, Switzerland
关键词
D O I
10.1162/089976600300015015
中图分类号
TP18 [人工智能理论];
学科分类号
081104 ; 0812 ; 0835 ; 1405 ;
摘要
Long short-term memory (LSTM; Hochreiter & Schmidhuber, 1997) can solve numerous tasks not solvable by previous learning algorithms for recurrent neural networks (RNNs). We identify a weakness of LSTM networks processing continual input streams that are not a priori segmented into subsequences with explicitly marked ends at which the network's internal state could be reset. Without resets, the state may grow indefinitely and eventually cause the network to break down. Our remedy is a novel, adaptive "forget gate" that enables an LSTM cell to learn to reset itself at appropriate times, thus releasing internal resources. We review illustrative benchmark problems on which standard LSTM outperforms other RNN algorithms. All algorithms (including LSTM) fail to solve continual versions of these problems. LSTM with forget gates, however, easily solves them, and in an elegant way.
引用
收藏
页码:2451 / 2471
页数:21
相关论文
共 22 条
  • [1] [Anonymous], 1999, IDSIA0199
  • [2] [Anonymous], 1991, Advances in Neural Information Processing Systems
  • [3] LEARNING LONG-TERM DEPENDENCIES WITH GRADIENT DESCENT IS DIFFICULT
    BENGIO, Y
    SIMARD, P
    FRASCONI, P
    [J]. IEEE TRANSACTIONS ON NEURAL NETWORKS, 1994, 5 (02): : 157 - 166
  • [4] Finite State Automata and Simple Recurrent Networks
    Cleeremans, Axel
    Servan-Schreiber, David
    McClelland, James L.
    [J]. NEURAL COMPUTATION, 1989, 1 (03) : 372 - 381
  • [5] Cummins F, 1999, P EUROSPEECH 99, P371
  • [6] DARKEN CJ, 1995, HDB BRAIN THEORY NEU, P941
  • [7] ADAPTIVE NEURAL OSCILLATOR USING CONTINUOUS-TIME BACK-PROPAGATION LEARNING
    DOYA, K
    YOSHIZAWA, S
    [J]. NEURAL NETWORKS, 1989, 2 (05) : 375 - 385
  • [8] Hochreiter S, 1997, NEURAL COMPUT, V9, P1735, DOI [10.1162/neco.1997.9.1.1, 10.1007/978-3-642-24797-2]
  • [9] Hochreiter S., 1991, Untersuchungen zu dynamischen neuronalen netzen
  • [10] Jordan M.I, 1986, P 8 ANN C COGN SCI S