← Search

Kuikui Liu

1 accepted papers

2024

Practical Performance Guarantees for Pipelined DNN Inference

ICML 2024spotlight

We optimize pipeline parallelism for deep neural network (DNN) inference by partitioning model graphs into $k$ stages and minimizing the running time of the bottleneck stage, including communication. We give practical and effective algorithms for this NP-hard problem, but our emphasis is on tackling…

Cited by 0SourcePDFScholar