2026
Beyond Classification Accuracy: Neural-MedBench and the Need for Deeper Reasoning Benchmarks
ICLR 2026poster
Recent advances in vision-language models (VLMs) have achieved remarkable performance on standard medical benchmarks, yet their true clinical reasoning ability remains unclear. Existing datasets predominantly emphasize classification accuracy, creating an evaluation illusion in which models appear p…