← Search

Jonni Kanerva

1 accepted papers

2021

Sparse is Enough in Scaling Transformers

NeurIPS 2021poster

Large Transformer models yield impressive results on many tasks, but are expensive to train, or even fine-tune, and so slow at decoding that their use and study becomes out of reach. We address this problem by leveraging sparsity. We study sparse variants for all layers in the Transformer and propos…

Cited by 101SourcePDFScholar