LSTM: A Search Space Odyssey
Presents the first large-scale analysis of eight LSTM variants on three tasks, finding the forget gate and output activation most critical.
Since the LSTM's 1995 inception, many variants have become state-of-the-art, raising interest in which components matter. This paper reports the first large-scale comparison of eight LSTM variants on speech recognition, handwriting recognition, and polyphonic music modeling. Hyperparameters were tuned per task by random search and ranked by functional ANOVA, over 5400 runs (~15 years of CPU time). No variant significantly beats the standard LSTM; the forget gate and output activation are its most critical components, and its hyperparameters are largely independent.
Based on: LSTM: A Search Space Odyssey · IEEE Transactions on Neural Networks and Learning Systems