2024
SciBench: Evaluating College-Level Scientific Problem-Solving Abilities of Large Language Models
ICML 2024poster
Most existing Large Language Model (LLM) benchmarks on scientific problem reasoning focus on problems grounded in high-school subjects and are confined to elementary algebraic operations. To systematically examine the reasoning capabilities required for solving complex scientific problems, we introd…