Efficient Hybrid Language Model Compression through Group-Aware SSM Pruning
Ali Taghibakhshi, Sharath Turuvekere Sreenivas, Saurav Muralidharan, Marcin Chochowski, Yashaswi Karnati, Raviraj Bhuminand Joshi, Ameya Sunil Mahabaleshwarkar, ZIJIA CHEN
Abstract
Hybrid language models that combine Attention and State Space Models (SSMs) have been shown to achieve state-of-the-art accuracy and runtime performance. Recent work has also demonstrated that applying pruning and distillation to Attention-only models yields smaller, more accurate models at a fraction of the training cost. In this work, we explore the effectiveness of compressing Hybrid architectures. To this end, we introduce a novel group-aware pruning method for Mamba layers that preserves the structural integrity of SSM blocks and their sequence modeling capabilities. We combine this method with FFN, embedding dimension, and layer pruning, along with knowledge distillation-based retraining to obtain a unified compression recipe for hybrid models. Using this recipe, we compress the Nemotron-H 8B Hybrid model down to 4B parameters with up to $40\times$ fewer training tokens compared to similarly-sized models. The resulting model surpasses the accuracy of similarly-sized models while achieving $\sim2\times$ faster inference throughput, significantly advancing the Pareto frontier.
BibTeX
@inproceedings{
taghibakhshi2025efficient,
title={Efficient Hybrid Language Model Compression through Group-Aware {SSM} Pruning},
author={Ali Taghibakhshi and Sharath Turuvekere Sreenivas and Saurav Muralidharan and Marcin Chochowski and Yashaswi Karnati and Raviraj Bhuminand Joshi and Ameya Sunil Mahabaleshwarkar and ZIJIA CHEN and Yoshi Suhara and Oluwatobi Olabiyi and Daniel Korzekwa and Mostofa Patwary and Mohammad Shoeybi and Jan Kautz and Bryan Catanzaro and Ashwath Aithal and Nima Tajbakhsh and Pavlo Molchanov},
booktitle={The Thirty-ninth Annual Conference on Neural Information Processing Systems},
year={2025},
url={https://openreview.net/forum?id=m3huAdsaGI}
}