← Search

Yijin Guo

1 accepted papers

2026

MedOmni-45°: A Safety–Performance Benchmark for Reasoning-Oriented LLMs in Medicine

AAAI 2026technical

With the rapid integration of large language models (LLMs) into medical decision-support aids, ensuring reliability in reasoning steps—not just final answers—is increasingly critical. Two key safety dimensions are Chain-of-Thought (CoT) faithfulness, which assesses alignment of the model’s reasoning

Cited by 0SourcePDFScholar