2026
MedOmni-45°: A Safety–Performance Benchmark for Reasoning-Oriented LLMs in Medicine
AAAI 2026technical
With the rapid integration of large language models (LLMs) into medical decision-support aids, ensuring reliability in reasoning steps—not just final answers—is increasingly critical. Two key safety dimensions are Chain-of-Thought (CoT) faithfulness, which assesses alignment of the model’s reasoning