← Search

Tobias Schnabel

8 accepted papers

2026

A Course Correction in Steerability Evaluation: Revealing Miscalibration and Side Effects in LLMs

AAAI 2026technical

Despite advances in large language models (LLMs) on reasoning and instruction-following benchmarks, it is unclear whether they can reliably produce outputs aligned with a variety of user goals, a concept called steerability. We highlight two gaps in current LLM evaluations for assessing steerability

Cited by 0SourcePDFScholar
2026

Reasoning about Reasoning: BAPO Bounds on Chain-of-Thought Token Complexity in LLMs

ICML 2026poster

Inference-time scaling via chain-of-thought (CoT) reasoning is a major driver of state-of-the-art LLM performance, but it comes with substantial latency and compute costs. We address a fundamental theoretical question: *how many* reasoning tokens are required to solve a problem as input size grows? …

Cited by 0SourceScholar
2025

Lost in Transmission: When and Why LLMs Fail to Reason Globally

NeurIPS 2025spotlight

Despite their many successes, transformer-based large language models (LLMs) continue to struggle with tasks that require complex reasoning over large parts of their input. We argue that these failures arise due to capacity limits on the accurate flow of information within LLMs. To formalize this is…

Cited by 0SourceScholar
2024

On Overcoming Miscalibrated Conversational Priors in LLM-based ChatBots

UAI 2024poster

We explore the use of Large Language Model (LLM-based) chatbots to power recommender systems. We observe that the chatbots respond poorly when they encounter under-specified requests (e.g., they make incorrect assumptions, hedge with a long response, or refuse to answer). We conjecture that such mi…

Cited by 4SourcePDFScholar
2024

Symbolic Prompt Program Search: A Structure-Aware Approach to Efficient Compile-Time Prompt Optimization

EMNLP 2024finding

In many modern LLM applications, such as retrieval augmented generation, prompts have become programs themselves. In these settings, prompt programs are repeatedly called with different user queries or data instances. A big practical challenge is optimizing such prompt programs. Recent work has most…

2021

Keep It Simple: Unsupervised Simplification of Multi-Paragraph Text

ACL 2021long

This work presents Keep it Simple (KiS), a new approach to unsupervised text simplification which learns to balance a reward across three properties: fluency, salience and simplicity. We train the model with a novel algorithm to optimize the reward (k-SCST), in which the model proposes several candi…

2019

Deep Generalized Method of Moments for Instrumental Variable Analysis

NeurIPS 2019poster

Instrumental variable analysis is a powerful tool for estimating causal effects when randomization or full control of confounders is not possible. The application of standard methods such as 2SLS, GMM, and more recent variants are significantly impeded when the causal effects are complex, the instru…

2016

Recommendations as Treatments: Debiasing Learning and Evaluation

ICML 2016poster

Most data for evaluating and training recommender systems is subject to selection biases, either through self-selection by the users or through the actions of the recommendation system itself. In this paper, we provide a principled approach to handle selection biases by adapting models and estimatio…

Cited by 834SourcePDFScholar