← Search

Jeff Da

7 accepted papers

2026

Imitation Learning for Multi-turn LM Agents via On-policy Expert Corrections

ICML 2026poster

A popular paradigm for training LM agents relies on imitation learning, fine-tuning on expert trajectories. However, we show that the off-policy nature of imitation learning for multi-turn LM agents suffers from the fundamental limitation known as covariate shift: as the student policy's behavior di…

Cited by 0SourceScholar
2026

SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?

ICML 2026poster

We present SWE-Bench Pro, a comprehensive benchmark designed to evaluate software engineering capabilities through complex, realistic programming challenges. This benchmark extends beyond traditional algorithmic problems to encompass the full spectrum of professional software development tasks. The …

Cited by 0SourceScholar
2024

A Careful Examination of Large Language Model Performance on Grade School Arithmetic

NeurIPS 2024spotlight

Large language models (LLMs) have achieved impressive success on many benchmarks for mathematical reasoning. However, there is growing concern that some of this performance actually reflects dataset contamination, where data closely resembling benchmark questions leaks into the training data, instea…

Cited by 77SourcePDFScholar
2024

Learning Goal-Conditioned Representations for Language Reward Models

NeurIPS 2024poster

Techniques that learn improved representations via offline data or self-supervised objectives have shown impressive results in traditional reinforcement learning. Nevertheless, it is unclear how improved representation learning can benefit reinforcement learning from human feedback on language model…

2023

Do Androids Laugh at Electric Sheep? Humor “Understanding” Benchmarks from The New Yorker Caption Contest

ACL 2023long

Large neural networks can now generate jokes, but do they really “understand” humor? We challenge AI models with three tasks derived from the New Yorker Cartoon Caption Contest: matching a joke to a cartoon, identifying a winning caption, and explaining why a winning caption is funny. These tasks en…

2021

(Comet-) Atomic 2020: On Symbolic and Neural Commonsense Knowledge Graphs

AAAI 2021technical

Recent years have brought about a renewed interest in commonsense representation and reasoning in the field of natural language understanding. The development of new commonsense knowledge graphs (CSKG) has been central to these advances as their diverse facts can be used and referenced by machine le…

2021

Edited Media Understanding Frames: Reasoning About the Intent and Implications of Visual Misinformation

ACL 2021long

Understanding manipulated media, from automatically generated ‘deepfakes’ to manually edited ones, raises novel research challenges. Because the vast majority of edited or manipulated images are benign, such as photoshopped images for visual enhancements, the key challenge is to understand the compl…

Cited by 17SourcePDFScholar