2024
AutoChunk: Automated Activation Chunk for Memory-Efficient Deep Learning Inference
ICLR 2024poster
Large deep learning models have achieved impressive performance across a range of applications. However, their large memory requirements, including parameter memory and activation memory, have become a significant challenge for their practical serving. While existing methods mainly address parameter…