2025
SciAssess: Benchmarking LLM Proficiency in Scientific Literature Analysis
NAACL 2025findings
Recent breakthroughs in Large Language Models (LLMs) have revolutionized scientific literature analysis. However, existing benchmarks fail to adequately evaluate the proficiency of LLMs in this domain, particularly in scenarios requiring higher-level abilities beyond mere memorization and the handli…