2026
MHB: Medical Hallucination Benchmark for Large Language Models in Complex Clinical Tasks
AAAI 2026technical
The integration of Large Language Models (LLMs) into clinical applications presents transformative potential but is undermined by the critical risk of hallucination, the generation of plausible but factually incorrect information. Such failures pose a direct threat to patient safety and the integrit