2017
Frustratingly Short Attention Spans in Neural Language Modeling
ICLR 2017poster
Current language modeling architectures often use recurrent neural networks. Recently, various methods for incorporating differentiable memory into these architectures have been proposed. When predicting the next token, these models query information from a memory of the recent history and thus can…