← Search

Karl Krauth

5 accepted papers

2026

Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces

ICLR 2026poster

AI agents may soon become capable of autonomously completing valuable, long-horizon tasks in diverse domains. Current benchmarks either do not measure real-world tasks, or are not sufficiently difficult to meaningfully measure frontier models. To this end, we present Terminal-Bench 1.5: a carefully…

Cited by 0SourcecodeScholar
2023

Modeling content creator incentives on algorithm-curated platforms

ICLR 2023top-5%

Content creators compete for user attention. Their reach crucially depends on algorithmic choices made by developers on online platforms. To maximize exposure, many creators adapt strategically, as evidenced by examples like the sprawling search engine optimization industry. This begets competition…

Cited by 44SourcePDFScholar
2021

On Component Interactions in Two-Stage Recommender Systems

NeurIPS 2021poster

Thanks to their scalability, two-stage recommenders are used by many of today's largest online platforms, including YouTube, LinkedIn, and Pinterest. These systems produce recommendations in two steps: (i) multiple nominators—tuned for low prediction latency—preselect a small subset of candidates fr…

Cited by 39SourcePDFScholar
2020

The Effect of Natural Distribution Shift on Question Answering Models

ICML 2020poster

We build four new test sets for the Stanford Question Answering Dataset (SQuAD) and evaluate the ability of question-answering systems to generalize to new data. Our first test set is from the original Wikipedia domain and measures the extent to which existing systems overfit the original test set.…

Cited by 186SourcePDFScholar
2019

Finite-time Analysis of Approximate Policy Iteration for the Linear Quadratic Regulator

NeurIPS 2019poster

We study the sample complexity of approximate policy iteration (PI) for the Linear Quadratic Regulator (LQR), building on a recent line of work using LQR as a testbed to understand the limits of reinforcement learning (RL) algorithms on continuous control tasks. Our analysis quantifies the tension b…

Cited by 73SourcePDFScholar