← Search

Haanju Yoo

2 accepted papers

2026

From Conversation to Query Execution: Benchmarking User and Tool Interactions for EHR Database Agents

ICLR 2026poster

Despite the impressive performance of LLM-powered agents, their adoption for Electronic Health Record (EHR) data access remains limited by the absence of benchmarks that adequately capture real-world clinical data access flows. In practice, two core challenges hinder deployment: query ambiguity from…

Cited by 0SourcecodeScholar
2026

Position: The Open Benchmark Paradox Must Be Resolved through Sovereign Medical Evaluation

ICML 2026poster

As medical large language models become increasingly involved in clinical actions, public benchmarks are often treated as proxies of deployment-readiness. However, this reliance creates a false sense of security because public scores are often based on data the models have already seen. We call this…

Cited by 0SourceScholar