2022
LiteTransformerSearch: Training-free Neural Architecture Search for Efficient Language Models
Mojan Javaheripi, Gustavo Henrique de Rosa, Subhabrata Mukherjee, Shital Shah, Tomasz Lukasz Religa, Caio Cesar Teodoro Mendes +3
NeurIPS 2022accept
The Transformer architecture is ubiquitously used as the building block of largescale autoregressive language models. However, finding architectures with the optimal trade-off between task performance (perplexity) and hardware constraints like peak memory utilization and latency is non-trivial. This…