← Search

Sumitra Ganesh

16 accepted papers

2025

Approximate Equivariance in Reinforcement Learning

AISTATS 2025poster

Equivariant neural networks have shown great success in reinforcement learning, improving sample efficiency and generalization when there is symmetry in the task. However, in many problems, only approximate symmetry is present, which makes imposing exact symmetry inappropriate. Recently, approximate…

Cited by 0SourcecodeScholar
2025

Collab: Controlled Decoding using Mixture of Agents for LLM Alignment

ICLR 2025poster

Alignment of Large Language models (LLMs) is crucial for safe and trustworthy deployment in applications. Reinforcement learning from human feedback (RLHF) has emerged as an effective technique to align LLMs to human preferences, and broader utilities, but it requires updating billions of model para…

Cited by 1SourcePDFScholar
2025

Decentralized Convergence to Equilibrium Prices in Trading Networks

AAAI 2025technical

We propose a decentralized market model in which agents can negotiate bilateral contracts. This builds on a similar, but centralized, model of trading networks introduced by Hatfield et al. in 2013. Prior work has established that fully-substitutable preferences guarantee the existence of competitiv…

Cited by 0SourcePDFScholar
2025

GenARM: Reward Guided Generation with Autoregressive Reward Model for Test-Time Alignment

ICLR 2025poster

Large Language Models (LLMs) exhibit impressive capabilities but require careful alignment with human preferences. Traditional training-time methods finetune LLMs using human preference datasets but incur significant training costs and require repeated training to handle diverse user preferences. Te…

2025

Learning in Herding Mean Field Games: Single-Loop Algorithm with Finite-Time Convergence Analysis

AISTATS 2025poster

We consider discrete-time stationary mean field games (MFG) with unknown dynamics and design algorithms for finding the equilibrium with finite-time complexity guarantees. Prior solutions to the problem assume either the contraction of a mean field optimality-consistency operator or strict weak mono…

Cited by 0SourceScholar
2025

Learning in Stackelberg Mean Field Games: A Non-Asymptotic Analysis

NeurIPS 2025poster

We study policy optimization in Stackelberg mean field games (MFGs), a hierarchical framework for modeling the strategic interaction between a single leader and an infinitely large population of homogeneous followers. The objective can be formulated as a structured bi-level optimization problem, in…

Cited by 0SourceScholar
2025

Prune 'n Predict: Optimizing LLM Decision-making with Conformal Prediction

ICML 2025poster

Large language models (LLMs) are empowering decision-making in several applications, including tool or API usage and answering multiple-choice questions (MCQs). However, incorrect outputs pose significant risks in high-stakes domains like healthcare and finance. To quantify LLM uncertainty and there…

Cited by 0SourcePDFScholar
2024

Efficient Inverse Multiagent Learning

ICLR 2024spotlight

In this paper, we study inverse game theory (resp. inverse multiagent learning) in which the goal is to find parameters of a game’s payoff functions for which the expected (resp. sampled) behavior is an equilibrium. We formulate these problems as generative-adversarial (i.e., min-max) optimization p…

Cited by 4SourcePDFScholar
2024

Information-Directed Pessimism for Offline Reinforcement Learning

ICML 2024poster

Policy optimization from batch data, i.e., offline reinforcement learning (RL) is important when collecting data from a current policy is not possible. This setting incurs distribution mismatch between batch training data and trajectories from the current policy. Pessimistic offsets estimate mismatc…

Cited by 1SourcePDFScholar
2023

Certifiably Robust Policy Learning against Adversarial Multi-Agent Communication

ICLR 2023poster

Communication is important in many multi-agent reinforcement learning (MARL) problems for agents to share information and make good decisions. However, when deploying trained communicative agents in a real-world application where noise and potential attackers exist, the safety of communication-based…

Cited by 21SourcePDFScholar
2022

Consensus Multiplicative Weights Update: Learning to Learn using Projector-based Game Signatures

ICML 2022spotlight

Cheung and Piliouras (2020) recently showed that two variants of the Multiplicative Weights Update method - OMWU and MWU - display opposite convergence properties depending on whether the game is zero-sum or cooperative. Inspired by this work and the recent literature on learning to optimize for sin…

Cited by 3SourcePDFScholar
2021

Factored Policy Gradients: Leveraging Structure for Efficient Learning in MOMDPs

NeurIPS 2021poster

Policy gradient methods can solve complex tasks but often fail when the dimensionality of the action-space or objective multiplicity grow very large. This occurs, in part, because the variance on score-based gradient estimators scales quadratically. In this paper, we address this problem through a f…

Cited by 10SourcePDFScholar
2020

Calibration of Shared Equilibria in General Sum Partially Observable Markov Games

NeurIPS 2020poster

Training multi-agent systems (MAS) to achieve realistic equilibria gives us a useful tool to understand and model real-world systems. We consider a general sum partially observable Markov game where agents of different types share a single policy network, conditioned on agent-specific information. T…

Cited by 17SourcePDFScholar