← Search

Sujay Bhatt

13 accepted papers

2025

Approximate Equivariance in Reinforcement Learning

AISTATS 2025poster

Equivariant neural networks have shown great success in reinforcement learning, improving sample efficiency and generalization when there is symmetry in the task. However, in many problems, only approximate symmetry is present, which makes imposing exact symmetry inappropriate. Recently, approximate…

Cited by 0SourcecodeScholar
2025

Collab: Controlled Decoding using Mixture of Agents for LLM Alignment

ICLR 2025poster

Alignment of Large Language models (LLMs) is crucial for safe and trustworthy deployment in applications. Reinforcement learning from human feedback (RLHF) has emerged as an effective technique to align LLMs to human preferences, and broader utilities, but it requires updating billions of model para…

Cited by 1SourcePDFScholar
2025

Decentralized Convergence to Equilibrium Prices in Trading Networks

AAAI 2025technical

We propose a decentralized market model in which agents can negotiate bilateral contracts. This builds on a similar, but centralized, model of trading networks introduced by Hatfield et al. in 2013. Prior work has established that fully-substitutable preferences guarantee the existence of competitiv…

Cited by 0SourcePDFScholar
2025

Learning in Herding Mean Field Games: Single-Loop Algorithm with Finite-Time Convergence Analysis

AISTATS 2025poster

We consider discrete-time stationary mean field games (MFG) with unknown dynamics and design algorithms for finding the equilibrium with finite-time complexity guarantees. Prior solutions to the problem assume either the contraction of a mean field optimality-consistency operator or strict weak mono…

Cited by 0SourceScholar
2025

Learning in Stackelberg Mean Field Games: A Non-Asymptotic Analysis

NeurIPS 2025poster

We study policy optimization in Stackelberg mean field games (MFGs), a hierarchical framework for modeling the strategic interaction between a single leader and an infinitely large population of homogeneous followers. The objective can be formulated as a structured bi-level optimization problem, in…

Cited by 0SourceScholar
2024

Information-Directed Pessimism for Offline Reinforcement Learning

ICML 2024poster

Policy optimization from batch data, i.e., offline reinforcement learning (RL) is important when collecting data from a current policy is not possible. This setting incurs distribution mismatch between batch training data and trajectories from the current policy. Pessimistic offsets estimate mismatc…

Cited by 1SourcePDFScholar
2023

Oracle-free Reinforcement Learning in Mean-Field Games along a Single Sample Path

AISTATS 2023poster

We consider online reinforcement learning in Mean-Field Games (MFGs). Unlike traditional approaches, we alleviate the need for a mean-field oracle by developing an algorithm that approximates the Mean-Field Equilibrium (MFE) using the single sample path of the generic agent. We call this Sandbox Lea…

Cited by 31SourcePDFScholar
2022

Minimax M-estimation under Adversarial Contamination

ICML 2022spotlight

We present a new finite-sample analysis of Catoni’s M-estimator under adversarial contamination, where an adversary is allowed to corrupt a fraction of the samples arbitrarily. We make minimal assumptions on the distribution of the uncontaminated random variables, namely, we only assume the existenc…

Cited by 8SourcePDFScholar
2022

Nearly Optimal Catoni’s M-estimator for Infinite Variance

ICML 2022spotlight

In this paper, we extend the remarkable M-estimator of Catoni \citep{Cat12} to situations where the variance is infinite. In particular, given a sequence of i.i.d random variables $\{X_i\}_{i=1}^n$ from distribution $\mathcal{D}$ over $\mathbb{R}$ with mean $\mu$, we only assume the existence of a k…

Cited by 17SourcePDFScholar