← Search

Ashutosh Nayyar

6 accepted papers

2025

Markov Balance Satisfaction Improves Performance in Strictly Batch Offline Imitation Learning

AAAI 2025technical

Imitation learning (IL) is notably effective for robotic tasks where directly programming behaviors or defining optimal control costs is challenging. In this work, we address a scenario where the imitator relies solely on observed behavior and cannot make environmental interactions during learning.…

2024

A Bayesian Learning Algorithm for Unknown Zero-sum Stochastic Games with an Arbitrary Opponent

AISTATS 2024poster

In this paper, we propose Posterior Sampling Reinforcement Learning for Zero-sum Stochastic Games (PSRL-ZSG), the first online learning algorithm that achieves Bayesian regret bound of $\tilde\mathcal{O}(HS\sqrt{AT})$ in the infinite-horizon zero-sum stochastic games with average-reward criterion. H…

Cited by 1SourcePDFScholar
2022

Optimal control of partially observable Markov decision processes with finite linear temporal logic constraints

UAI 2022poster

Autonomous agents often operate in environments where the state is partially observed. In addition to maximizing their cumulative reward, agents must execute complex tasks with rich temporal and logical structures. These tasks can be expressed using temporal logic languages like finite linear tempo…

Cited by 8SourcePDFScholar
2020

Regret Bounds for Decentralized Learning in Cooperative Multi-Agent Dynamical Systems

UAI 2020poster

Regret analysis is challenging in Multi-Agent Reinforcement Learning (MARL) primarily due to the dynamical environments and the decentralized information among agents. We attempt to solve this challenge in the context of decentralized learning in multi-agent linear-quadratic (LQ) dynamical systems.…

Cited by 12SourcePDFScholar
2017

Learning Unknown Markov Decision Processes: A Thompson Sampling Approach

NeurIPS 2017poster

We consider the problem of learning an unknown Markov Decision Process (MDP) that is weakly communicating in the infinite horizon setting. We propose a Thompson Sampling-based reinforcement learning algorithm with dynamic episodes (TSDE). At the beginning of each episode, the algorithm generates a s…

Cited by 167SourcePDFScholar