← Search

Adam Wierman

46 accepted papers

2026

Distributionally Robust Cooperative Multi-agent Reinforcement Learning with Value Factorization

ICLR 2026poster

Cooperative multi-agent reinforcement learning (MARL) commonly adopts centralized training with decentralized execution, where value-factorization methods enforce the individual-global-maximum (IGM) principle so that decentralized greedy actions recover the team-optimal joint action. However, the re…

Cited by 0SourceScholar
2025

Approximate Global Convergence of Independent Learning in Multi-Agent Systems

AISTATS 2025poster

Independent learning (IL) is a popular approach for achieving scalability in large-scale multi-agent systems, yet it typically lacks global convergence guarantees. In this paper, we study two representative algorithms—independent $Q$-learning and independent natural actor-critic—within both value-ba…

Cited by 0SourceScholar
2025

Breaking the Curse of Multiagency in Robust Multi-Agent Reinforcement Learning

ICML 2025poster

Standard multi-agent reinforcement learning (MARL) algorithms are vulnerable to sim-to-real gaps. To address this, distributionally robust Markov games (RMGs) have been proposed to enhance robustness in MARL by optimizing the worst-case performance when game dynamics shift within a prescribed uncert…

Cited by 5SourcePDFScholar
2025

Conformal Risk Training: End-to-End Optimization of Conformal Risk Control

NeurIPS 2025poster

While deep learning models often achieve high predictive accuracy, their predictions typically do not come with any provable guarantees on risk or reliability, which are critical for deployment in high-stakes applications. The framework of conformal risk control (CRC) provides a distribution-free, f…

Cited by 0SourceScholar
2025

Efficient Policy Optimization in Robust Constrained MDPs with Iteration Complexity Guarantees

NeurIPS 2025poster

Constrained decision-making is essential for designing safe policies in real-world control systems, yet simulated environments often fail to capture real-world adversities. We consider the problem of learning a policy that will maximize the cumulative reward while satisfying a constraint, even when…

Cited by 0SourceScholar
2025

Fusing Reward and Dueling Feedback in Stochastic Bandits

ICML 2025poster

This paper investigates the fusion of absolute (reward) and relative (dueling) feedback in stochastic bandits, where both feedback types are gathered in each decision round. We derive a regret lower bound, demonstrating that an efficient algorithm may incur only the smaller among the reward…

Cited by 0SourcePDFScholar
2025

Hybrid Transfer Reinforcement Learning: Provable Sample Efficiency from Shifted-Dynamics Data

AISTATS 2025oral

Online reinforcement learning (RL) typically requires online interaction data to learn a policy for a target task, but collecting such data can be high-stakes. This prompts interest in leveraging historical data to improve sample efficiency. The historical data may come from outdated or related sour…

Cited by 0SourcecodeScholar
2025

Maximizing the Value of Predictions in Control: Accuracy Is Not Enough

NeurIPS 2025poster

We study the value of stochastic predictions in online optimal control with random disturbances. Prior work provides performance guarantees based on prediction error but ignores the stochastic dependence between predictions and disturbances. We introduce a general framework modeling their joint dist…

Cited by 0SourcecodeScholar
2025

Online Robust Reinforcement Learning Through Monte-Carlo Planning

ICML 2025poster

Monte Carlo Tree Search (MCTS) is a powerful framework for solving complex decision-making problems, yet it often relies on the assumption that the simulator and the real-world dynamics are identical. Although this assumption helps achieve the success of MCTS in games like Chess, Go, and Shogi, the…

Cited by 0SourcePDFScholar
2025

Overcoming the Curse of Dimensionality in Reinforcement Learning Through Approximate Factorization

ICML 2025poster

Factored Markov Decision Processes (FMDPs) offer a promising framework for overcoming the curse of dimensionality in reinforcement learning (RL) by decomposing high-dimensional MDPs into smaller and independently evolving components. Despite their potential, existing studies on FMDPs face three key…

Cited by 1SourcePDFScholar
2025

Reinforcement Learning with Imperfect Transition Predictions: A Bellman-Jensen Approach

NeurIPS 2025spotlight

Traditional reinforcement learning (RL) assumes the agents make decisions based on Markov decision processes (MDPs) with one-step transition models. In many real-world applications, such as energy management and stock investment, agents can access multi-step predictions of future states, which provi…

Cited by 0SourceScholar
2025

Robust Gymnasium: A Unified Modular Benchmark for Robust Reinforcement Learning

ICLR 2025poster

Driven by inherent uncertainty and the sim-to-real gap, robust reinforcement learning (RL) seeks to improve resilience against the complexity and variability in agent-environment sequential interactions. Despite the existence of a large number of RL benchmarks, there is a lack of standardized benchm…

Cited by 1SourcePDFScholar
2025

SPiDR: A Simple Approach for Zero-Shot Safety in Sim-to-Real Transfer

NeurIPS 2025poster

Deploying reinforcement learning (RL) safely in the real world is challenging, as policies trained in simulators must face the inevitable *sim-to-real gap*. Robust safe RL techniques are provably safe, however difficult to scale, while domain randomization is more practical yet prone to unsafe behav…

Cited by 0SourceScholar
2024

Best of Both Worlds Guarantees for Smoothed Online Quadratic Optimization

ICML 2024poster

We study the smoothed online quadratic optimization (SOQO) problem where, at each round $t$, a player plays an action $x_t$ in response to a quadratic hitting cost and an additional squared $\ell_2$-norm cost for switching actions. This problem class has strong connections to a wide range of applica…

Cited by 1SourcePDFScholar
2024

Chasing Convex Functions with Long-term Constraints

ICML 2024poster

We introduce and study a family of online metric problems with long-term constraints. In these problems, an online player makes decisions $\mathbf{x}_t$ in a metric space $(X,d)$ to simultaneously minimize their hitting cost $f_t(\mathbf{x}_t)$ and switching cost as determined by the metric. Over th…

Cited by 3SourcePDFScholar
2024

Enhancing Efficiency of Safe Reinforcement Learning via Sample Manipulation

NeurIPS 2024poster

Safe reinforcement learning (RL) is crucial for deploying RL agents in real-world applications, as it aims to maximize long-term rewards while satisfying safety constraints. However, safe RL often suffers from sample inefficiency, requiring extensive interactions with the environment to learn a safe…

2024

Learning the Uncertainty Sets of Linear Control Systems via Set Membership: A Non-asymptotic Analysis

ICML 2024poster

This paper studies uncertainty set estimation for unknown linear systems. Uncertainty sets are crucial for the quality of robust control since they directly influence the conservativeness of the control design. Departing from the confidence region analysis of least squares estimation, this paper foc…

Cited by 3SourcePDFScholar
2024

Model-Free Robust $\phi$-Divergence Reinforcement Learning Using Both Offline and Online Data

ICML 2024poster

The robust $\phi$-regularized Markov Decision Process (RRMDP) framework focuses on designing control policies that are robust against parameter uncertainties due to mismatches between the simulator (nominal) model and real-world settings. This work makes *two* important contributions. First, we prop…

Cited by 5SourcePDFScholar
2024

Near-Optimal Distributionally Robust Reinforcement Learning with General $L_p$ Norms

NeurIPS 2024poster

To address the challenges of sim-to-real gap and sample efficiency in reinforcement learning (RL), this work studies distributionally robust Markov decision processes (RMDPs) --- optimize the worst-case performance when the deployed environment is within an uncertainty set around some nominal MDP. D…

Cited by 0SourcePDFScholar
2024

Online Algorithms with Uncertainty-Quantified Predictions

ICML 2024poster

The burgeoning field of algorithms with predictions studies the problem of using possibly imperfect machine learning predictions to improve online algorithm performance. While nearly all existing algorithms in this framework make no assumptions on prediction quality, a number of methods providing un…

Cited by 5SourcePDFScholar
2024

Sample-Efficient Robust Multi-Agent Reinforcement Learning in the Face of Environmental Uncertainty

ICML 2024poster

To overcome the sim-to-real gap in reinforcement learning (RL), learned policies must maintain robustness against environmental uncertainties. While robust RL has been widely studied in single-agent regimes, in multi-agent environments, the problem remains understudied---despite the fact that the pr…

Cited by 13SourcePDFScholar
2023

A Finite-Sample Analysis of Payoff-Based Independent Learning in Zero-Sum Stochastic Games

NeurIPS 2023poster

In this work, we study two-player zero-sum stochastic games and develop a variant of the smoothed best-response learning dynamics that combines independent learning dynamics for matrix games with the minimax value iteration for stochastic games. The resulting learning dynamics are payoff-based, conv…

Cited by 14SourcePDFScholar
2023

Adversarial Attacks on Online Learning to Rank with Click Feedback

NeurIPS 2023poster

Online learning to rank (OLTR) is a sequential decision-making problem where a learning agent selects an ordered list of items and receives feedback through user clicks. Although potential attacks against OLTR algorithms may cause serious losses in real-world applications, there is limited knowledge…

Cited by 6SourcePDFScholar
2023

Anytime-Competitive Reinforcement Learning with Policy Prior

NeurIPS 2023poster

This paper studies the problem of Anytime-Competitive Markov Decision Process (A-CMDP). Existing works on Constrained Markov Decision Processes (CMDPs) aim to optimize the expected reward while constraining the expected cost over random dynamics, but the cost in a specific episode can still be unsat…

Cited by 2SourcePDFScholar
2023

Beyond Black-Box Advice: Learning-Augmented Algorithms for MDPs with Q-Value Predictions

NeurIPS 2023poster

We study the tradeoff between consistency and robustness in the context of a single-trajectory time-varying Markov Decision Process (MDP) with untrusted machine-learned advice. Our work departs from the typical approach of treating advice as coming from black-box sources by instead considering a set…

Cited by 4SourcePDFScholar
2023

Contextual Combinatorial Bandits with Probabilistically Triggered Arms

ICML 2023poster

We study contextual combinatorial bandits with probabilistically triggered arms (C$^2$MAB-T) under a variety of smoothness conditions that capture a wide range of applications, such as contextual cascading bandits and contextual influence maximization bandits. Under the triggering probability modula…

Cited by 21SourcePDFScholar
2023

Convergence rates for localized actor-critic in networked Markov potential games

UAI 2023poster

We introduce a class of networked Markov potential games where agents are associated with nodes in a network. Each agent has its own local potential function, and the reward of each agent depends only on the states and actions of agents within a neighborhood. In this context, we propose a localized…

2023

Online Adaptive Policy Selection in Time-Varying Systems: No-Regret via Contractive Perturbations

NeurIPS 2023poster

We study online adaptive policy selection in systems with time-varying costs and dynamics. We develop the Gradient-based Adaptive Policy Selection (GAPS) algorithm together with a general analytical framework for online policy selection via online optimization. Under our proposed notion of contracti…

Cited by 16SourcePDFScholar
2023

Optimal robustness-consistency tradeoffs for learning-augmented metrical task systems

AISTATS 2023poster

We examine the problem of designing learning-augmented algorithms for metrical task systems (MTS) that exploit machine-learned advice while maintaining rigorous, worst-case guarantees on performance. We propose an algorithm, DART, that achieves this dual objective, providing cost within a multiplica…

Cited by 36SourcePDFScholar
2023

Robust Learning for Smoothed Online Convex Optimization with Feedback Delay

NeurIPS 2023poster

We study a general form of Smoothed Online Convex Optimization, a.k.a. SOCO, including multi-step switching costs and feedback delay. We propose a novel machine learning (ML) augmented online algorithm, Robustness-Constrained Learning (RCL), which combines untrusted ML predictions with a trusted exp…

Cited by 4SourcePDFScholar
2023

SustainGym: Reinforcement Learning Environments for Sustainable Energy Systems

NeurIPS 2023poster

The lack of standardized benchmarks for reinforcement learning (RL) in sustainability applications has made it difficult to both track progress on specific domains and identify bottlenecks for researchers to focus their efforts. In this paper, we present SustainGym, a suite of five environments desi…

2022

Bounded-Regret MPC via Perturbation Analysis: Prediction Error, Constraints, and Nonlinearity

NeurIPS 2022accept

We study Model Predictive Control (MPC) and propose a general analysis pipeline to bound its dynamic regret. The pipeline first requires deriving a perturbation bound for a finite-time optimal control problem. Then, the perturbation bound is used to bound the per-step error of MPC, which leads to a…

Cited by 16SourcePDFScholar
2022

Decentralized Online Convex Optimization in Networked Systems

ICML 2022spotlight

We study the problem of networked online convex optimization, where each agent individually decides on an action at every time step and agents cooperatively seek to minimize the total global cost over a finite horizon. The global cost is made up of three types of local costs: convex node costs, temp…

Cited by 11SourcePDFScholar
2021

Data-driven Competitive Algorithms for Online Knapsack and Set Cover

AAAI 2021technical

The design of online algorithms has tended to focus on algorithms with worst-case guarantees, e.g., bounds on the competitive ratio. However, it is well-known that such algorithms are often overly pessimistic, performing sub-optimally on non-worst-case inputs. In this paper, we develop an approach…

Cited by 39SourcePDFScholar
2021

Multi-Agent Reinforcement Learning in Stochastic Networked Systems

NeurIPS 2021poster

We study multi-agent reinforcement learning (MARL) in a stochastic network of agents. The objective is to find localized policies that maximize the (discounted) global reward. In general, scalability is a challenge in this setting because the size of the global state/action space can be exponential…

Cited by 51SourcePDFScholar
2021

Pareto-Optimal Learning-Augmented Algorithms for Online Conversion Problems

NeurIPS 2021poster

This paper leverages machine-learned predictions to design competitive algorithms for online conversion problems with the goal of improving the competitive ratio when predictions are accurate (i.e., consistency), while also guaranteeing a worst-case competitive ratio regardless of the prediction qua…

Cited by 39SourcePDFScholar
2021

Perturbation-based Regret Analysis of Predictive Control in Linear Time Varying Systems

NeurIPS 2021spotlight

We study predictive control in a setting where the dynamics are time-varying and linear, and the costs are time-varying and well-conditioned. At each time step, the controller receives the exact predictions of costs, dynamics, and disturbances for the future $k$ time steps. We show that when the pre…

Cited by 46SourcePDFScholar
2020

Online Optimization with Memory and Competitive Control

NeurIPS 2020poster

This paper presents competitive algorithms for a novel class of online optimization problems with memory. We consider a setting where the learner seeks to minimize the sum of a hitting cost and a switching cost that depends on the previous $p$ decisions. This setting generalizes Smoothed Online Conv…

Cited by 63SourcePDFScholar
2020

Scalable Multi-Agent Reinforcement Learning for Networked Systems with Average Reward

NeurIPS 2020poster

It has long been recognized that multi-agent reinforcement learning (MARL) faces significant scalability issues due to the fact that the size of the state and action spaces are exponentially large in the number of agents. In this paper, we identify a rich class of networked MARL problems where the m…

Cited by 92SourcePDFScholar
2019

Beyond Online Balanced Descent: An Optimal Algorithm for Smoothed Online Optimization

NeurIPS 2019spotlight

We study online convex optimization in a setting where the learner seeks to minimize the sum of a per-round hitting cost and a movement cost which is incurred when changing decisions between rounds. We prove a new lower bound on the competitive ratio of any online algorithm in the setting where the…

Cited by 79SourcePDFScholar