2025
MMSciBench: Benchmarking Language Models on Chinese Multimodal Scientific Problems
ACL 2025finding
Recent advances in large language models (LLMs) and vision-language models (LVLMs) have shown promise across many tasks, yet their scientific reasoning capabilities remain untested, particularly in multimodal settings. We present MMSciBench, a benchmark for evaluating mathematical and physical reaso…