← Search

Guangyang LU

1 accepted papers

2024

AutoChunk: Automated Activation Chunk for Memory-Efficient Deep Learning Inference

ICLR 2024poster

Large deep learning models have achieved impressive performance across a range of applications. However, their large memory requirements, including parameter memory and activation memory, have become a significant challenge for their practical serving. While existing methods mainly address parameter…

Cited by 0SourcePDFScholar