← Search

Heegyu Kim

3 accepted papers

2026

DEER: A Benchmark for Evaluating Deep Research Agents on Expert Report Generation

ICML 2026poster

Recent advances in large language models have enabled deep research systems that generate expert-level reports through multi-step reasoning and evidence-based synthesis. However, evaluating such reports remains challenging: report quality is multifaceted, making it difficult to determine what to ass…

Cited by 0SourceScholar
2025

FLEX: Expert-level False-Less EXecution Metric for Text-to-SQL Benchmark

NAACL 2025long

Text-to-SQL systems have become crucial for translating natural language into SQL queries in various industries, enabling non-technical users to perform complex data operations. The need for accurate evaluation methods has increased as these systems have grown more sophisticated. However, the Execut…