NAACL 2025findings28 citations

SciAssess: Benchmarking LLM Proficiency in Scientific Literature Analysis

Hengxing Cai, Xiaochen Cai, Junhan Chang, Sihang Li, Lin Yao, Wang Changxin, Zhifeng Gao, Hongshuai Wang

Abstract

Recent breakthroughs in Large Language Models (LLMs) have revolutionized scientific literature analysis. However, existing benchmarks fail to adequately evaluate the proficiency of LLMs in this domain, particularly in scenarios requiring higher-level abilities beyond mere memorization and the handling of multimodal data.In response to this gap, we introduce SciAssess, a benchmark specifically designed for the comprehensive evaluation of LLMs in scientific literature analysis. It aims to thoroughly assess the efficacy of LLMs by evaluating their capabilities in Memorization (L1), Comprehension (L2), and Analysis & Reasoning (L3). It encompasses a variety of tasks drawn from diverse scientific fields, including biology, chemistry, material, and medicine.To ensure the reliability of SciAssess, rigorous quality control measures have been implemented, ensuring accuracy, anonymization, and compliance with copyright standards. SciAssess evaluates 11 LLMs, highlighting their strengths and areas for improvement. We hope this evaluation supports the ongoing development of LLM applications in scientific literature analysis.SciAssess and its resources are available at https://github.com/sci-assess/SciAssess.

BibTeX
@inproceedings{cai-etal-2025-sciassess,
    title = "{S}ci{A}ssess: Benchmarking {LLM} Proficiency in Scientific Literature Analysis",
    author = "Cai, Hengxing  and
      Cai, Xiaochen  and
      Chang, Junhan  and
      Li, Sihang  and
      Yao, Lin  and
      Changxin, Wang  and
      Gao, Zhifeng  and
      Wang, Hongshuai  and
      Yongge, Li  and
      Lin, Mujie  and
      Yang, Shuwen  and
      Wang, Jiankun  and
      Xu, Mingjun  and
      Huang, Jin  and
      Fang, Xi  and
      Zhuang, Jiaxi  and
      Yin, Yuqi  and
      Li, Yaqi  and
      Chen, Changhong  and
      Cheng, Zheng  and
      Zhao, Zifeng  and
      Zhang, Linfeng  and
      Ke, Guolin",
    editor = "Chiruzzo, Luis  and
      Ritter, Alan  and
      Wang, Lu",
    booktitle = "Findings of the Association for Computational Linguistics: NAACL 2025",
    month = apr,
    year = "2025",
    address = "Albuquerque, New Mexico",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2025.findings-naacl.125/",
    pages = "2335--2357",
    ISBN = "979-8-89176-195-7"
}
SciAssess: Benchmarking LLM Proficiency in Scientific Literature Analysis · NAACL 2025