← Search

Lucas Jun Koba Sato

3 accepted papers

2025

Measuring AI Ability to Complete Long Software Tasks

NeurIPS 2025poster

Despite rapid progress on AI benchmarks, the real-world meaning of benchmark performance remains unclear. To quantify the capabilities of AI systems in terms of human capabilities, we propose a new metric: 50%-task-completion time horizon. This is the time humans typically take to complete tasks tha…

Cited by 0SourceScholar
2025

RE-Bench: Evaluating Frontier AI R&D Capabilities of Language Model Agents against Human Experts

ICML 2025spotlight

Frontier AI safety policies highlight automation of AI research and development (R&D) by AI agents as an important capability to anticipate. However, there exist few evaluations for AI R&D capabilities, and none that are highly realistic and have a direct comparison to human performance. We introduc…

Cited by 16SourcePDFScholar
2022

A Few-Shot Semantic Parser for Wizard-of-Oz Dialogues with the Precise ThingTalk Representation

ACL 2022findings

Previous attempts to build effective semantic parsers for Wizard-of-Oz (WOZ) conversations suffer from the difficulty in acquiring a high-quality, manually annotated training set. Approaches based only on dialogue synthesis are insufficient, as dialogues generated from state-machine based models are…

Cited by 10SourcePDFScholar