← Search

David Wentzlaff

1 accepted papers

2024

Kraken: Inherently Parallel Transformers For Efficient Multi-Device Inference

NeurIPS 2024poster

Large Transformer networks are increasingly used in settings where low inference latency is necessary to enable new applications and improve the end-user experience. However, autoregressive inference is resource intensive and requires parallelism for efficiency. Parallelism introduces collective com…

Cited by 2SourcePDFScholar