← Search

Eugene Wu

2 accepted papers

2026

LakeQA: A Benchmark for Complex Exploratory QA over a Million-Scale Data Lake

ICML 2026poster

Recent large language models (LLMs) have shown rapid progress on reading-based question answering (QA), where the evidence is explicitly provided or trivially retrievable. In contrast, real-world questions are often not paired with accurate evidence documents. The useful evidence resides in a massiv…

Cited by 0SourceScholar
2026

Outrunning LLM Cutoffs: A Live Kernel Crash Resolution Benchmark for All

ICML 2026poster

Repairing system crashes discovered by kernel fuzzers like Syzkaller is a critical yet underexplored challenge in software engineering. While recent works have introduced Large Language Model (LLM) based agents for Linux kernel crash-resolution, their evaluation benchmarks are usually static and thu…

Cited by 0SourceScholar