← Search

Stephen Mcaleer

7 accepted papers

2024

Automated Design of Affine Maximizer Mechanisms in Dynamic Settings

AAAI 2024technical

Dynamic mechanism design is a challenging extension to ordinary mechanism design in which the mechanism designer must make a sequence of decisions over time in the face of possibly untruthful reports of participating agents. Optimizing dynamic mechanisms for welfare is relatively well understood. Ho…

Cited by 9SourcePDFScholar
2024

Policy Space Response Oracles: A Survey

IJCAI 2024poster

Game theory provides a mathematical way to study the interaction between multiple decision makers. However, classical game-theoretic analysis is limited in scalability due to the large number of strategies, precluding direct application to more complex scenarios. This survey provides a comprehensive…

Cited by 14SourcePDFScholar
2024

Scalable Mechanism Design for Multi-Agent Path Finding

IJCAI 2024poster

Multi-Agent Path Finding (MAPF) involves determining paths for multiple agents to travel simultaneously and collision-free through a shared area toward given goal locations. This problem is computationally complex, especially when dealing with large numbers of agents, as is common in realistic appli…

2022

Reducing Variance in Temporal-Difference Value Estimation via Ensemble of Deep Networks

ICML 2022spotlight

In temporal-difference reinforcement learning algorithms, variance in value estimation can cause instability and overestimation of the maximal target value. Many algorithms have been proposed to reduce overestimation, including several recent ensemble methods, however none have shown success in samp…

2020

Evolutionary Reinforcement Learning for Sample-Efficient Multiagent Coordination

ICML 2020poster

Many cooperative multiagent reinforcement learning environments provide agents with a sparse team-based reward, as well as a dense agent-specific reward that incentivizes learning basic skills. Training policies solely on the team-based reward is often difficult due to its sparsity. Also, relying so…

Cited by 79SourcePDFScholar
2020

Pipeline PSRO: A Scalable Approach for Finding Approximate Nash Equilibria in Large Games

NeurIPS 2020poster

Finding approximate Nash equilibria in zero-sum imperfect-information games is challenging when the number of information states is large. Policy Space Response Oracles (PSRO) is a deep reinforcement learning algorithm grounded in game theory that is guaranteed to converge to an approximate Nash equ…

2019

Solving the Rubik's Cube with Approximate Policy Iteration

ICLR 2019poster

Recently, Approximate Policy Iteration (API) algorithms have achieved super-human proficiency in two-player zero-sum games such as Go, Chess, and Shogi without human data. These API algorithms iterate between two policies: a slow policy (tree search), and a fast policy (a neural network). In these t…

Cited by 52SourcePDFScholar