2024
Kraken: Inherently Parallel Transformers For Efficient Multi-Device Inference
NeurIPS 2024poster
Large Transformer networks are increasingly used in settings where low inference latency is necessary to enable new applications and improve the end-user experience. However, autoregressive inference is resource intensive and requires parallelism for efficiency. Parallelism introduces collective com…