← Search

Ruishan Liu

4 accepted papers

2026

Cancer-Myth: Evaluating Large Language Models on Patient Questions with False Presuppositions

ICLR 2026poster

Cancer patients are increasingly turning to large language models (LLMs) for medical information, making it critical to assess how well these models handle complex, personalized questions. However, current medical benchmarks focus on medical exams or consumer-searched questions and do not evaluate…

Cited by 0SourcecodeScholar
2026

CounselBench: A Large-Scale Expert Evaluation and Adversarial Benchmarking of Large Language Models in Mental Health Question Answering

ICLR 2026oral

Medical question answering (QA) benchmarks often focus on multiple-choice or fact-based tasks, leaving open-ended answers to real patient questions underexplored. This gap is particularly critical in mental health, where patient questions often mix symptoms, treatment concerns, and emotional needs,…

Cited by 0SourcecodeScholar