← Search

Suhas Kamasetty Ramesh

1 accepted papers

2025

On the Generalization vs Fidelity Paradox in Knowledge Distillation

ACL 2025finding

Knowledge distillation (KD) is a key technique for compressing large language models into smaller ones while preserving performance. Despite the recent traction of KD research, its effectiveness for smaller language models (LMs) and the mechanisms driving knowledge transfer remain underexplored. In…