← Search

Emanuel Ben Baruch

2 accepted papers

2025

Distilling the Knowledge in Data Pruning

ICML 2025poster

With the increasing size of datasets used for training neural networks, data pruning has gained traction in recent years. However, most current data pruning algorithms are limited in their ability to preserve accuracy compared to models trained on the full data, especially in high pruning regimes. I…

Cited by 5SourcePDFScholar
2025

Group-Aware Reinforcement Learning for Output Diversity in Large Language Models

EMNLP 2025

Large Language Models (LLMs) often suffer from mode collapse, repeatedly generating the same few completions even when many valid answers exist, limiting their diversity across a wide range of tasks. We introduce Group-Aware Policy Optimization (GAPO) , a simple extension of the recent and popular G