ACL 2025long0 citations

A Survey on Efficient Large Language Model Training: From Data-centric Perspectives

Junyu Luo, Bohan Wu, Xiao Luo, Zhiping Xiao, Yiqiao Jin, Rong-Cheng Tu, Nan Yin, Yifan Wang

Abstract

Post-training of Large Language Models (LLMs) is crucial for unlocking their task generalization potential and domain-specific capabilities. However, the current LLM post-training paradigm faces significant data challenges, including the high costs of manual annotation and diminishing marginal returns on data scales. Therefore, achieving data-efficient post-training has become a key research question. In this paper, we present the first systematic survey of data-efficient LLM post-training from a data-centric perspective. We propose a taxonomy of data-efficient LLM post-training methods, covering data selection, data quality enhancement, synthetic data generation, data distillation and compression, and self-evolving data ecosystems. We summarize representative approaches in each category and outline future research directions. By examining the challenges in data-efficient LLM post-training, we highlight open problems and propose potential research avenues. We hope our work inspires further exploration into maximizing the potential of data utilization in large-scale model training. Paper List: https://github.com/luo-junyu/Awesome-Data-Efficient-LLM

BibTeX
@inproceedings{luo-etal-2025-survey,
    title = "A Survey on Efficient Large Language Model Training: From Data-centric Perspectives",
    author = "Luo, Junyu  and
      Wu, Bohan  and
      Luo, Xiao  and
      Xiao, Zhiping  and
      Jin, Yiqiao  and
      Tu, Rong-Cheng  and
      Yin, Nan  and
      Wang, Yifan  and
      Yuan, Jingyang  and
      Ju, Wei  and
      Zhang, Ming",
    editor = "Che, Wanxiang  and
      Nabende, Joyce  and
      Shutova, Ekaterina  and
      Pilehvar, Mohammad Taher",
    booktitle = "Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = jul,
    year = "2025",
    address = "Vienna, Austria",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2025.acl-long.1493/",
    doi = "10.18653/v1/2025.acl-long.1493",
    pages = "30904--30920",
    ISBN = "979-8-89176-251-0"
}
A Survey on Efficient Large Language Model Training: From Data-centric Perspectives · ACL 2025