← Search

Srinivas Shakkottai

12 accepted papers

2025

CONGO: Compressive Online Gradient Optimization

ICLR 2025poster

We address the challenge of zeroth-order online convex optimization where the objective function's gradient exhibits sparsity, indicating that only a small number of dimensions possess non-zero gradients. Our aim is to leverage this sparsity to obtain useful estimates of the objective function's gra…

Cited by 0SourcePDFScholar
2025

DOPL: Direct Online Preference Learning for Restless Bandits with Preference Feedback

ICLR 2025poster

Restless multi-armed bandits (RMAB) has been widely used to model constrained sequential decision making problems, where the state of each restless arm evolves according to a Markov chain and each state transition generates a scalar reward. However, the success of RMAB crucially relies on the availa…

Cited by 2SourcePDFScholar
2025

Transformers are Provably Optimal In-context Estimators for Wireless Communications

AISTATS 2025poster

Pre-trained transformers exhibit the capability of adapting to new tasks through in-context learning (ICL), where they efficiently utilize a limited set of prompts without explicit model optimization. The canonical communication problem of estimating transmitted symbols from received observations c…

Cited by 0SourcecodeScholar
2024

Federated Ensemble-Directed Offline Reinforcement Learning

NeurIPS 2024poster

We consider the problem of federated offline reinforcement learning (RL), a scenario under which distributed learning agents must collaboratively learn a high-quality control policy only using small pre-collected datasets generated according to different unknown behavior policies. Na\"{i}vely combin…

2024

Risk-Averse Fine-tuning of Large Language Models

NeurIPS 2024poster

We consider the challenge of mitigating the generation of negative or toxic content by the Large Language Models (LLMs) in response to certain prompts. We propose integrating risk-averse principles into LLM fine-tuning to minimize the occurrence of harmful outputs, particularly rare but significant…

2022

DOPE: Doubly Optimistic and Pessimistic Exploration for Safe Reinforcement Learning

NeurIPS 2022accept

Safe reinforcement learning is extremely challenging--not only must the agent explore an unknown environment, it must do so while ensuring no safety constraint violations. We formulate this safe reinforcement learning (RL) problem using the framework of a finite-horizon Constrained Markov Decision…

2022

Enhanced Meta Reinforcement Learning via Demonstrations in Sparse Reward Environments

NeurIPS 2022accept

Meta reinforcement learning (Meta-RL) is an approach wherein the experience gained from solving a variety of tasks is distilled into a meta-policy. The meta-policy, when adapted over only a small (or just a single) number of steps, is able to perform near-optimally on a new, related task. However,…

Cited by 2SourcePDFScholar
2022

Reinforcement Learning with Sparse Rewards using Guidance from Offline Demonstration

ICLR 2022spotlight

A major challenge in real-world reinforcement learning (RL) is the sparsity of reward feedback. Often, what is available is an intuitive but sparse reward function that only indicates whether the task is completed partially or fully. However, the lack of carefully designed, fine grain feedback imp…

2021

Learning with Safety Constraints: Sample Complexity of Reinforcement Learning for Constrained MDPs

AAAI 2021technical

Many physical systems have underlying safety considerations that require that the policy employed ensures the satisfaction of a set of constraints. The analytical formulation usually takes the form of a Constrained Markov Decision Process (CMDP). We focus on the case where the CMDP is unknown, and…

Cited by 52SourcePDFScholar
2021

Model-Based Reinforcement Learning for Infinite-Horizon Discounted Constrained Markov Decision Processes

IJCAI 2021poster

In many real-world reinforcement learning (RL) problems, in addition to maximizing the objective, the learning agent has to maintain some necessary safety constraints. We formulate the problem of learning a safe policy as an infinite-horizon discounted Constrained Markov Decision Process (CMDP)…

Cited by 16SourcePDFScholar
2021

NeurWIN: Neural Whittle Index Network For Restless Bandits Via Deep RL

NeurIPS 2021poster

Whittle index policy is a powerful tool to obtain asymptotically optimal solutions for the notoriously intractable problem of restless bandits. However, finding the Whittle indices remains a difficult problem for many practical restless bandits with convoluted transition kernels. This paper proposes…

2021

Reinforcement Learning for Mean Field Games with Strategic Complementarities

AISTATS 2021poster

Mean Field Games (MFG) are the class of games with a very large number of agents and the standard equilibrium concept is a Mean Field Equilibrium (MFE). Algorithms for learning MFE in dynamic MFGs are unknown in general. Our focus is on an important subclass that possess a monotonicity property call…

Cited by 18SourcePDFScholar