← Search

Thomas Kwa

4 accepted papers

2025

Measuring AI Ability to Complete Long Software Tasks

NeurIPS 2025poster

Despite rapid progress on AI benchmarks, the real-world meaning of benchmark performance remains unclear. To quantify the capabilities of AI systems in terms of human capabilities, we propose a new metric: 50%-task-completion time horizon. This is the time humans typically take to complete tasks tha…

Cited by 0SourceScholar
2024

Catastrophic Goodhart: regularizing RLHF with KL divergence does not mitigate heavy-tailed reward misspecification

NeurIPS 2024poster

When applying reinforcement learning from human feedback (RLHF), the reward is learned from data and, therefore, always has some error. It is common to mitigate this by regularizing the policy with KL divergence from a base model, with the hope that balancing reward with regularization will achieve…

2024

Compact Proofs of Model Performance via Mechanistic Interpretability

NeurIPS 2024poster

We propose using mechanistic interpretability -- techniques for reverse engineering model weights into human-interpretable algorithms -- to derive and compactly prove formal guarantees on model performance. We prototype this approach by formally proving accuracy lower bounds for a small transformer…

Cited by 5SourcePDFScholar
2024

InterpBench: Semi-Synthetic Transformers for Evaluating Mechanistic Interpretability Techniques

NeurIPS 2024poster

Mechanistic interpretability methods aim to identify the algorithm a neural network implements, but it is difficult to validate such methods when the true algorithm is unknown. This work presents InterpBench, a collection of semi-synthetic yet realistic transformers with known circuits for evaluatin…

Cited by 4SourcePDFScholar