Importance-Aware Data Selection for Efficient LLM Instruction Tuning
Tingyu Jiang, Shen Li, Yiyao Song, Lan Zhang, Hualei Zhu, Yuan Zhao, Xiaohang Xu, Kenjiro Taura
Abstract
Instruction tuning plays a critical role in enhancing the performance and efficiency of Large Language Models (LLMs). Its success depends not only on the quality of the instruction data but also on the inherent capabilities of the LLM itself. Some studies suggest that even a small amount of high-quality data can achieve instruction fine-tuning results that are on par with, or even exceed, those from using a full-scale dataset. However, rather than focusing solely on calculating data quality scores to evaluate instruction data, there is a growing need to select high-quality data that maximally enhances the performance of instruction tuning for a given LLM. In this paper, we propose the Model Instruction Weakness Value (MIWV) as a novel metric to quantify the importance of instruction data in enhancing model
BibTeX
@inproceedings{aaai2026_importanceawared,
title = {Importance-Aware Data Selection for Efficient LLM Instruction Tuning},
author = {Tingyu Jiang and Shen Li and Yiyao Song and Lan Zhang and Hualei Zhu and Yuan Zhao and Xiaohang Xu and Kenjiro Taura and Hao Henry Wang},
booktitle = {AAAI 2026},
year = {2026}
}