2025
CS-Bench: A Comprehensive Benchmark for Large Language Models towards Computer Science Mastery
ICLR 2025poster
Large language models (LLMs) have demonstrated significant potential in advancing various fields of research and society. However, the current community of LLMs overly focuses on benchmarks for analyzing specific foundational skills (e.g. mathematics and code generation), neglecting an all-round eva…