← Search

Nathan Kallus

65 accepted papers

2026

ConvRec-R1: Training LLM-based Conversational Recommender Systems with Reinforcement Learning

ICLR 2026poster

Large language models (LLMs) are reshaping the recommender system paradigm by enabling users to express preferences and receive recommendations through conversations. Yet, aligning LLMs to the recommendation task remains challenging: pretrained LLMs often generate out-of-catalog items, violate requi…

Cited by 0SourcecodeScholar
2026

GAAVI: Global Asymptotic Anytime Valid Inference for the Conditional Mean Function

ICML 2026poster

Inference on the conditional mean function (CMF) is central to tasks from adaptive experimentation to optimal treatment assignment and algorithmic fairness auditing. In this work, we provide a novel asymptotic anytime-valid test for a CMF global null (e.g., that all conditional means are zero) and c…

Cited by 0SourceScholar
2025

$Q\sharp$: Provably Optimal Distributional RL for LLM Post-Training

NeurIPS 2025poster

Reinforcement learning (RL) post-training is crucial for LLM alignment and reasoning, but existing policy-based methods, such as PPO and DPO, can fall short of fixing shortcuts inherited from pre-training. In this work, we introduce $Q\sharp$, a value-based algorithm for KL-regularized RL that guide…

Cited by 0SourcecodeScholar
2025

A Reductions Approach to Risk-Sensitive Reinforcement Learning with Optimized Certainty Equivalents

ICML 2025poster

We study risk-sensitive RL where the goal is learn a history-dependent policy that optimizes some risk measure of cumulative rewards. We consider a family of risks called the optimized certainty equivalents (OCE), which captures important risk measures such as conditional value-at-risk (CVaR), entro…

Cited by 0SourcePDFScholar
2025

GST-UNet: A Neural Framework for Spatiotemporal Causal Inference with Time-Varying Confounding

NeurIPS 2025poster

Estimating causal effects from spatiotemporal observational data is essential in public health, environmental science, and policy evaluation, where randomized experiments are often infeasible. Existing approaches, however, either rely on strong structural assumptions or fail to handle key challenges…

Cited by 0SourceScholar
2025

LLM-based Conversational Recommendation Agents with Collaborative Verbalized Experience

EMNLP 2025

Large language models (LLMs) have demonstrated impressive zero-shot capabilities in conversational recommender systems (CRS). However, effectively utilizing historical conversations remains a significant challenge. Current approaches either retrieve few-shot examples or extract global rules to enhan

2025

Multi-Armed Bandits with Interference: Bridging Causal Inference and Adversarial Bandits

ICML 2025poster

Experimentation with interference poses a significant challenge in contemporary online platforms. Prior research on experimentation with interference has concentrated on the final output of a policy. Cumulative performance, while equally important, is less well understood. To address this gap, we in…

Cited by 0SourcePDFScholar
2025

Reward Maximization for Pure Exploration: Minimax Optimal Good Arm Identification for Nonparametric Multi-Armed Bandits

AISTATS 2025poster

In multi-armed bandits, reward maximization and pure exploration are often at odds with each other. The former focuses on exploiting arms with the highest means, while the latter may require constant exploration across all arms. In this work, we focus on good arm identification (GAI), a pure explora…

Cited by 0SourceScholar
2025

Value-Guided Search for Efficient Chain-of-Thought Reasoning

NeurIPS 2025poster

In this paper, we propose a simple and efficient method for value model training on long-context reasoning traces. Compared to existing process reward models (PRMs), our method does not require a fine-grained notion of ``step,'' which is difficult to define for long-context reasoning models. By coll…

Cited by 0SourcecodeScholar
2024

Efficient and Sharp Off-Policy Evaluation in Robust Markov Decision Processes

NeurIPS 2024poster

We study the evaluation of a policy under best- and worst-case perturbations to a Markov decision process (MDP), using transition observations from the original MDP, whether they are generated under the same or a different policy. This is an important problem when there is the possibility of a shift…

2024

Estimating Heterogeneous Treatment Effects by Combining Weak Instruments and Observational Data

NeurIPS 2024poster

Accurately predicting conditional average treatment effects (CATEs) is crucial in personalized medicine and digital platform analytics. Since the treatments of interest often cannot be directly randomized, observational data is leveraged to learn CATEs, but this approach can incur significant bias…

2024

Inferring the Long-Term Causal Effects of Long-Term Treatments from Short-Term Experiments

ICML 2024oral

We study inference on the long-term causal effect of a continual exposure to a novel intervention, which we term a long-term treatment, based on an experiment involving only short-term observations. Key examples include the long-term health effects of regularly-taken medicine or of environmental haz…

2024

More Benefits of Being Distributional: Second-Order Bounds for Reinforcement Learning

ICML 2024poster

In this paper, we prove that Distributional Reinforcement Learning (DistRL), which learns the return distribution, can obtain second-order bounds in both online and offline RL in general settings with function approximation. Second-order bounds are instance-dependent bounds that scale with the varia…

Cited by 15SourcePDFScholar
2024

Peeking with PEAK: Sequential, Nonparametric Composite Hypothesis Tests for Means of Multiple Data Streams

ICML 2024poster

We propose a novel nonparametric sequential test for composite hypotheses for means of multiple data streams. Our proposed method, peeking with expectation-based averaged capital (PEAK), builds upon the testing-by-betting framework and provides a non-asymptotic $\alpha$-level test across any stoppin…

2024

Provable Offline Preference-Based Reinforcement Learning

ICLR 2024spotlight

In this paper, we investigate the problem of offline Preference-based Reinforcement Learning (PbRL) with human feedback where feedback is available in the form of preference between trajectory pairs rather than explicit rewards. Our proposed algorithm consists of two main steps: (1) estimate the imp…

Cited by 41SourcePDFScholar
2024

Switching the Loss Reduces the Cost in Batch Reinforcement Learning

ICML 2024poster

We propose training fitted Q-iteration with log-loss (FQI-LOG) for batch reinforcement learning (RL). We show that the number of samples needed to learn a near-optimal policy with FQI-LOG scales with the accumulated cost of the optimal policy, which is zero in problems where acting optimally achieve…

Cited by 6SourcePDFScholar
2023

B-Learner: Quasi-Oracle Bounds on Heterogeneous Causal Effects Under Hidden Confounding

ICML 2023poster

Estimating heterogeneous treatment effects from observational data is a crucial task across many fields, helping policy and decision-makers take better actions. There has been recent progress on robust and efficient methods for estimating the conditional average treatment effect (CATE) function, but…

Cited by 27SourcePDFScholar
2023

Computationally Efficient PAC RL in POMDPs with Latent Determinism and Conditional Embeddings

ICML 2023poster

We study reinforcement learning with function approximation for large-scale Partially Observable Markov Decision Processes (POMDPs) where the state space and observation space are large or even continuous. Particularly, we consider Hilbert space embeddings of POMDP where the feature of latent states…

Cited by 14SourcePDFScholar
2023

Future-Dependent Value-Based Off-Policy Evaluation in POMDPs

NeurIPS 2023spotlight

We study off-policy evaluation (OPE) for partially observable MDPs (POMDPs) with general function approximation. Existing methods such as sequential importance sampling estimators and fitted-Q evaluation suffer from the curse of horizon in POMDPs. To circumvent this problem, we develop a novel model…

2023

Offline Minimax Soft-Q-learning Under Realizability and Partial Coverage

NeurIPS 2023poster

We consider offline reinforcement learning (RL) where we only have only access to offline data. In contrast to numerous offline RL algorithms that necessitate the uniform coverage of the offline data over state and action space, we propose value-based algorithms with PAC guarantees under partial cov…

Cited by 8SourcePDFScholar
2023

Provable Safe Reinforcement Learning with Binary Feedback

AISTATS 2023poster

Safety is a crucial necessity in many applications of reinforcement learning (RL), whether robotic, automotive, or medical. Many existing approaches to safe RL rely on receiving numeric safety feedback, but in many cases this feedback can only take binary values; that is, whether an action in a give…

2023

Robust and Agnostic Learning of Conditional Distributional Treatment Effects

AISTATS 2023poster

The conditional average treatment effect (CATE) is the best measure of individual causal effects given baseline covariates. However, the CATE only captures the (conditional) average, and can overlook risks and tail events, which are important to treatment choice. In aggregate analyses, this is usual…

2023

The Benefits of Being Distributional: Small-Loss Bounds for Reinforcement Learning

NeurIPS 2023poster

While distributional reinforcement learning (DistRL) has been empirically effective, the question of when and why it is better than vanilla, non-distributional RL has remained unanswered. This paper explains the benefits of DistRL through the lens of small-loss bounds, which are instance-dependent b…

2022

Doubly Robust Distributionally Robust Off-Policy Evaluation and Learning

ICML 2022spotlight

Off-policy evaluation and learning (OPE/L) use offline observational data to make better decisions, which is crucial in applications where online experimentation is limited. However, depending entirely on logged data, OPE/L is sensitive to environment distribution shifts — discrepancies between the…

2022

Learning Bellman Complete Representations for Offline Policy Evaluation

ICML 2022oral

We study representation learning for Offline Reinforcement Learning (RL), focusing on the important task of Offline Policy Evaluation (OPE). Recent work shows that, in contrast to supervised learning, realizability of the Q-function is not enough for learning it. Two sufficient conditions for sample…

2022

Provably Efficient Reinforcement Learning in Partially Observable Dynamical Systems

NeurIPS 2022accept

We study Reinforcement Learning for partially observable systems using function approximation. We propose a new PO-bilinear framework, that is general enough to include models such as undercomplete tabular Partially Observable Markov Decision Processes (POMDPs), Linear Quadratic Gaussian (LQG), Pred…

Cited by 41SourcePDFScholar
2021

Control Variates for Slate Off-Policy Evaluation

NeurIPS 2021poster

We study the problem of off-policy evaluation from batched contextual bandit data with multidimensional actions, often termed slates. The problem is common to recommender systems and user-interface optimization, and it is particularly challenging because of the combinatorially-sized action space. Sw…

2021

Off-policy Evaluation in Infinite-Horizon Reinforcement Learning with Latent Confounders

AISTATS 2021poster

Off-policy evaluation (OPE) in reinforcement learning is an important problem in settings where experimentation is limited, such as healthcare. But, in these very same settings, observed actions are often confounded by unobserved variables making OPE even more difficult. We study an OPE problem in a…

Cited by 56SourcePDFScholar
2021

Post-Contextual-Bandit Inference

NeurIPS 2021poster

Contextual bandit algorithms are increasingly replacing non-adaptive A/B tests in e-commerce, healthcare, and policymaking because they can both improve outcomes for study participants and increase the chance of identifying good or even best policies. To support credible inference on novel intervent…

Cited by 54SourcePDFScholar
2021

Risk Minimization from Adaptively Collected Data: Guarantees for Supervised and Policy Learning

NeurIPS 2021poster

Empirical risk minimization (ERM) is the workhorse of machine learning, whether for classification and regression or for off-policy policy learning, but its model-agnostic guarantees can fail when we use adaptively collected data, such as the result of running a contextual bandit algorithm. We study…

Cited by 17SourcePDFScholar
2020

DeepMatch: Balancing Deep Covariate Representations for Causal Inference Using Adversarial Training

ICML 2020poster

We study optimal covariate balance for causal inferences from observational data when rich covariates and complex relationships necessitate flexible modeling with neural networks. Standard approaches such as propensity weighting and matching/balancing fail in such settings due to miscalibrated prope…

Cited by 95SourcePDFScholar
2020

Doubly Robust Off-Policy Value and Gradient Estimation for Deterministic Policies

NeurIPS 2020poster

Offline reinforcement learning, wherein one uses off-policy data logged by a fixed behavior policy to evaluate and learn new policies, is crucial in applications where experimentation is limited such as medicine. We study the estimation of policy value and gradient of a deterministic policy from off…

Cited by 16SourcePDFScholar
2019

Assessing Disparate Impact of Personalized Interventions: Identifiability and Bounds

NeurIPS 2019poster

Personalized interventions in social services, education, and healthcare leverage individual-level causal effect predictions in order to give the best treatment to each individual or to prioritize program interventions for the individuals most likely to benefit. While the sensitivity of these domain…

2019

Deep Generalized Method of Moments for Instrumental Variable Analysis

NeurIPS 2019poster

Instrumental variable analysis is a powerful tool for estimating causal effects when randomization or full control of confounders is not possible. The application of standard methods such as 2SLS, GMM, and more recent variants are significantly impeded when the causal effects are complex, the instru…

2019

Interval Estimation of Individual-Level Causal Effects Under Unobserved Confounding

AISTATS 2019poster

We study the problem of learning conditional average treatment effects (CATE) from observational data with unobserved confounders. The CATE function maps baseline covariates to individual causal effect predictions and is key for personalized assessments. Recent work has focused on how to learn CATE…

Cited by 122SourcePDFScholar
2019

Intrinsically Efficient, Stable, and Bounded Off-Policy Evaluation for Reinforcement Learning

NeurIPS 2019poster

Off-policy evaluation (OPE) in both contextual bandits and reinforcement learning allows one to evaluate novel decision policies without needing to conduct exploration, which is often costly or otherwise infeasible. The problem's importance has attracted many proposed solutions, including importance…

Cited by 60SourcePDFScholar
2019

The Fairness of Risk Scores Beyond Classification: Bipartite Ranking and the XAUC Metric

NeurIPS 2019poster

Where machine-learned predictive risk scores inform high-stakes decisions, such as bail and sentencing in criminal justice, fairness has been a serious concern. Recent work has characterized the disparate impact that such risk scores can have when used for a binary classification task. This may not…

2018

Causal Inference with Noisy and Missing Covariates via Matrix Factorization

NeurIPS 2018poster

Valid causal inference in observational studies often requires controlling for confounders. However, in practice measurements of confounders may be noisy, and can lead to biased estimates of causal effects. We show that we can reduce bias induced by measurement noise using a large number of noisy me…