← Search

Asuman E. Ozdaglar

26 accepted papers

2026

Beyond RLHF and NLHF: Population-Proportional Alignment under an Axiomatic Framework

ICLR 2026poster

Conventional preference learning methods often prioritize opinions held more widely when aggregating preferences from multiple evaluators. This may result in policies that are biased in favor of some types of opinions or groups and susceptible to strategic manipulation. To address this issue, we de…

Cited by 0SourceScholar
2026

Online Learning and Equilibrium Computation with Ranking Feedback

ICLR 2026oral

Online learning in arbitrary and possibly adversarial environments has been extensively studied in sequential decision-making, with a strong connection to equilibrium computation in game theory. Most existing online learning algorithms are based on \emph{numeric} utility feedback from the environmen…

Cited by 0SourceScholar
2025

A Policy-Gradient Approach to Solving Imperfect-Information Games with Best-Iterate Convergence

ICLR 2025poster

Policy gradient methods have become a staple of any single-agent reinforcement learning toolbox, due to their combination of desirable properties: iterate convergence, efficient use of stochastic trajectory feedback, and theoretically-sound avoidance of importance sampling corrections. In multi-agen…

Cited by 3SourcePDFScholar
2025

Contextual Optimization Under Model Misspecification: A Tractable and Generalizable Approach

ICML 2025poster

Contextual optimization problems are prevalent in decision-making applications where historical data and contextual features are used to learn predictive models that inform optimal actions. However, practical applications often suffer from model misspecification due to incomplete knowledge of the un…

Cited by 0SourcePDFScholar
2025

Do LLM Agents Have Regret? A Case Study in Online Learning and Games

ICLR 2025poster

Large language models (LLMs) have been increasingly employed for (interactive) decision-making, via the development of LLM-based autonomous agents. Despite their emerging successes, the performance of LLM agents in decision-making has not been fully investigated through quantitative metrics, especia…

Cited by 20SourcePDFScholar
2025

MAPoRL: Multi-Agent Post-Co-Training for Collaborative Large Language Models with Reinforcement Learning

ACL 2025long

Leveraging multi-agentic frameworks to enhance large language models (LLMs) has demonstrated significant potential recently, with most existing studies focusing on prompting and developing workflows with frozen LLMs. In this paper, we aim to further unleash the power of such multi-agentic frameworks…

Cited by 0SourcePDFScholar
2025

What Data Enables Optimal Decisions? An Exact Characterization for Linear Optimization

NeurIPS 2025poster

We study the fundamental question of how informative a dataset is for solving a given decision-making task. In our setting, the dataset provides partial information about unknown parameters that influence task outcomes. Focusing on linear programs, we characterize when a dataset is sufficient to rec…

Cited by 0SourceScholar
2024

A Unified Linear Programming Framework for Offline Reward Learning from Human Demonstrations and Feedback

ICML 2024poster

Inverse Reinforcement Learning (IRL) and Reinforcement Learning from Human Feedback (RLHF) are pivotal methodologies in reward learning, which involve inferring and shaping the underlying reward function of sequential decision-making problems based on observed human demonstrations and feedback. Most…

Cited by 1SourcePDFScholar
2024

MisinfoEval: Generative AI in the Era of “Alternative Facts”

EMNLP 2024main

The spread of misinformation on social media platforms threatens democratic processes, contributes to massive economic losses, and endangers public health. Many efforts to address misinformation focus on a knowledge deficit model and propose interventions for improving users’ critical thinking throu…

Cited by 3SourcePDFScholar
2024

Uniformly Stable Algorithms for Adversarial Training and Beyond

ICML 2024poster

In adversarial machine learning, neural networks suffer from a significant issue known as robust overfitting, where the robust test accuracy decreases over epochs (Rice et al., 2020). Recent research conducted by Xing et al., 2021;Xiao et al., 2022 has focused on studying the uniform stability of ad…

2023

A Finite-Sample Analysis of Payoff-Based Independent Learning in Zero-Sum Stochastic Games

NeurIPS 2023poster

In this work, we study two-player zero-sum stochastic games and develop a variant of the smoothed best-response learning dynamics that combines independent learning dynamics for matrix games with the minimax value iteration for stochastic games. The resulting learning dynamics are payoff-based, conv…

Cited by 14SourcePDFScholar
2023

Multi-Player Zero-Sum Markov Games with Networked Separable Interactions

NeurIPS 2023poster

We study a new class of Markov games, \textit{(multi-player) zero-sum Markov Games} with {\it Networked separable interactions} (zero-sum NMGs), to model the local interaction structure in non-cooperative multi-agent sequential decision-making. We define a zero-sum NMG as a model where {the payoffs…

Cited by 11SourcePDFScholar
2023

Revisiting the Linear-Programming Framework for Offline RL with General Function Approximation

ICML 2023poster

Offline reinforcement learning (RL) aims to find an optimal policy for sequential decision-making using a pre-collected dataset, without further interaction with the environment. Recent theoretical progress has focused on developing sample-efficient offline RL algorithms with various relaxed assumpt…

Cited by 27SourcePDFScholar
2023

The Power of Regularization in Solving Extensive-Form Games

ICLR 2023poster

In this paper, we investigate the power of {\it regularization}, a common technique in reinforcement learning and optimization, in solving extensive-form games (EFGs). We propose a series of new algorithms based on regularizing the payoff functions of the game, and establish a set of convergence re…

Cited by 26SourcePDFScholar
2023

Time-Reversed Dissipation Induces Duality Between Minimizing Gradient Norm and Function Value

NeurIPS 2023poster

In convex optimization, first-order optimization methods efficiently minimizing function values have been a central subject study since Nesterov's seminal work of 1983. Recently, however, Kim and Fessler's OGM-G and Lee et al.'s FISTA-G have been presented as alternatives that efficiently minimize t…

Cited by 17SourcePDFScholar
2022

Bridging Central and Local Differential Privacy in Data Acquisition Mechanisms

NeurIPS 2022accept

We study the design of optimal Bayesian data acquisition mechanisms for a platform interested in estimating the mean of a distribution by collecting data from privacy-conscious users. In our setting, users have heterogeneous sensitivities for two types of privacy losses corresponding to local and ce…

Cited by 5SourcePDFScholar
2022

What is a Good Metric to Study Generalization of Minimax Learners?

NeurIPS 2022accept

Minimax optimization has served as the backbone of many machine learning problems. Although the convergence behavior of optimization algorithms has been extensively studied in minimax settings, their generalization guarantees, i.e., how the model trained on empirical data performs on the unseen test…

Cited by 16SourcePDFScholar
2021

Decentralized Q-learning in Zero-sum Markov Games

NeurIPS 2021poster

We study multi-agent reinforcement learning (MARL) in infinite-horizon discounted zero-sum Markov games. We focus on the practical but challenging setting of decentralized MARL, where agents make decisions without coordination by a centralized controller, but only based on their own payoffs and lo…

Cited by 121SourcePDFScholar
2021

Generalization of Model-Agnostic Meta-Learning Algorithms: Recurring and Unseen Tasks

NeurIPS 2021poster

In this paper, we study the generalization properties of Model-Agnostic Meta-Learning (MAML) algorithms for supervised learning problems. We focus on the setting in which we train the MAML model over $m$ tasks, each with $n$ data points, and characterize its generalization error from two points of v…

Cited by 66SourcePDFScholar
2021

On the Convergence Theory of Debiased Model-Agnostic Meta-Reinforcement Learning

NeurIPS 2021poster

We consider Model-Agnostic Meta-Learning (MAML) methods for Reinforcement Learning (RL) problems, where the goal is to find a policy using data from several tasks represented by Markov Decision Processes (MDPs) that can be updated by one step of \textit{stochastic} policy gradient for the realized M…

2019

Community Inference from Graph Signals with Hidden Nodes

ICASSP 2019accepted

Many recent works on inference of graph structure assume that the graph signals are fully observable. For large graphs with thousands or millions of nodes, this entails high complexity on the data collection and processing steps. Here, we study a community inference problem on partially observed (su…

Cited by 0SourceScholar
2018

Community Detection from Low-Rank Excitations of a Graph Filter

ICASSP 2018accepted

This paper considers the problem of inferring the topology of a graph from noisy outputs of an unknown graph filter excited by low-rank signals. Limited by this low-rank structure, we focus on solving the community detection problem, whose aim is to partition the node set of the unknown graph into s…

Cited by 0SourceScholar
2018

Identifying Susceptible Agents in Time Varying Opinion Dynamics Through Compressive Measurements

ICASSP 2018accepted

We provide a compressive-measurement based method to detect susceptible agents who may receive misinformation through their contact with `stubborn agents' whose goal is to influence the opinions of agents in the network. We consider a DeGroot-type opinion dynamics model where regular agents revise t…

Cited by 0SourceScholar