← Search

Heeyoung Kwak

5 accepted papers

2026

From Conversation to Query Execution: Benchmarking User and Tool Interactions for EHR Database Agents

ICLR 2026poster

Despite the impressive performance of LLM-powered agents, their adoption for Electronic Health Record (EHR) data access remains limited by the absence of benchmarks that adequately capture real-world clinical data access flows. In practice, two core challenges hinder deployment: query ambiguity from…

Cited by 0SourcecodeScholar
2026

Position: The Open Benchmark Paradox Must Be Resolved through Sovereign Medical Evaluation

ICML 2026poster

As medical large language models become increasingly involved in clinical actions, public benchmarks are often treated as proxies of deployment-readiness. However, this reliance creates a false sense of security because public scores are often based on data the models have already seen. We call this…

Cited by 0SourceScholar
2024

EHRNoteQA: An LLM Benchmark for Real-World Clinical Practice Using Discharge Summaries

NeurIPS 2024poster

Discharge summaries in Electronic Health Records (EHRs) are crucial for clinical decision-making, but their length and complexity make information extraction challenging, especially when dealing with accumulated summaries across multiple patient admissions. Large Language Models (LLMs) show promise…

2023

BREAK: Breaking the Dialogue State Tracking Barrier with Beam Search and Re-ranking

ACL 2023long

Despite the recent advances in dialogue state tracking (DST), the joint goal accuracy (JGA) of the existing methods on MultiWOZ 2.1 still remains merely 60%. In our preliminary error analysis, we find that beam search produces a pool of candidates that is likely to include the correct dialogue state…

2022

Subgraph Representation Learning with Hard Negative Samples for Inductive Link Prediction

ICASSP 2022accepted

The inductive link prediction in knowledge graphs (KGs) is often addressed to induce logical rules that capture entity-independent relational semantics. Recent studies suggest graph representation learning to encode these logical rules within the local subgraph structures. With this approach, the mo…

Cited by 0SourceScholar