ACL 2025long0 citations

CFBench: A Comprehensive Constraints-Following Benchmark for LLMs

Tao Zhang, ChengLIn Zhu, Yanjun Shen, Wenjing Luo, Yan Zhang, Hao Liang, Fan Yang, Mingan Lin

Abstract

The adeptness of Large Language Models (LLMs) in comprehending and following natural language instructions is critical for their deployment in sophisticated real-world applications. Existing evaluations mainly focus on fragmented constraints or narrow scenarios, but they overlook the comprehensiveness and authenticity of constraints from the user’s perspective. To bridge this gap, we propose CFBench, a large-scale Chinese Comprehensive Constraints Following Benchmark for LLMs, featuring 1,000 curated samples that cover more than 200 real-life scenarios and over 50 NLP tasks. CFBench meticulously compiles constraints from real-world instructions and constructs an innovative systematic framework for constraint types, which includes 10 primary categories and over 25 subcategories, and ensures each constraint is seamlessly integrated within the instructions. To make certain that the evaluation of LLM outputs aligns with user perceptions, we propose an advanced methodology that integrates multi-dimensional assessment criteria with requirement prioritization, covering various perspectives of constraints, instructions, and requirement fulfillment. Evaluating current leading LLMs on CFBench reveals substantial room for improvement in constraints following, and we further investigate influencing factors and enhancement strategies. The data and code will be made available.

BibTeX
@inproceedings{zhang-etal-2025-cfbench,
    title = "{CFB}ench: A Comprehensive Constraints-Following Benchmark for {LLM}s",
    author = "Zhang, Tao  and
      Zhu, ChengLIn  and
      Shen, Yanjun  and
      Luo, Wenjing  and
      Zhang, Yan  and
      Liang, Hao  and
      Zhang, Tao  and
      Yang, Fan  and
      Lin, Mingan  and
      Qiao, Yujing  and
      Chen, Weipeng  and
      Cui, Bin  and
      Zhang, Wentao  and
      Zhou, Zenan",
    editor = "Che, Wanxiang  and
      Nabende, Joyce  and
      Shutova, Ekaterina  and
      Pilehvar, Mohammad Taher",
    booktitle = "Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = jul,
    year = "2025",
    address = "Vienna, Austria",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2025.acl-long.1581/",
    doi = "10.18653/v1/2025.acl-long.1581",
    pages = "32926--32944",
    ISBN = "979-8-89176-251-0"
}
CFBench: A Comprehensive Constraints-Following Benchmark for LLMs · ACL 2025