CQIL: Inference Latency Optimization with Concurrent Computation of Quasi-Independent Layers
The fast-growing large scale language models are delivering unprecedented performance on almost all natural language processing tasks. However, the effectiveness of large language models are reliant on an exponentially increasing number of parameters. The overwhelming computation complexity incurs a…