2024
CQIL: Inference Latency Optimization with Concurrent Computation of Quasi-Independent Layers
ACL 2024long
The fast-growing large scale language models are delivering unprecedented performance on almost all natural language processing tasks. However, the effectiveness of large language models are reliant on an exponentially increasing number of parameters. The overwhelming computation complexity incurs a…