← Search

Stephen M Mcaleer

2 accepted papers

2022

Independent Natural Policy Gradient always converges in Markov Potential Games

AISTATS 2022poster

Natural policy gradient has emerged as one of the most successful algorithms for computing optimal policies in challenging Reinforcement Learning (RL) tasks, yet, very little was known about its convergence properties until recently. The picture becomes more blurry when it comes to multi-agent RL (M…

Cited by 64SourcePDFScholar
2022

Proving Theorems using Incremental Learning and Hindsight Experience Replay

ICML 2022spotlight

Traditional automated theorem proving systems for first-order logic depend on speed-optimized search and many handcrafted heuristics designed to work over a wide range of domains. Machine learning approaches in the literature either depend on these traditional provers to bootstrap themselves, by lev…

Cited by 23SourcePDFScholar