← Search

Hanna Mazzawi

1 accepted papers

2024

Deep Fusion: Efficient Network Training via Pre-trained Initializations

ICML 2024poster

Training deep neural networks for large language models (LLMs) remains computationally very expensive. To mitigate this, network growing algorithms offer potential cost savings, but their underlying mechanisms are poorly understood. In this paper, we propose a theoretical framework using backward er…

Cited by 4SourcePDFScholar