← Search

Akshat Kumar

18 accepted papers

2026

Revisiting Distribution Correction Estimation for Offline Imitation Learning with Suboptimal Dataset

ICML 2026poster

Imitation Learning (IL) has demonstrated strong capabilities in learning high-quality policies from expert demonstrations for sequential decision-making tasks. Nonetheless, its effectiveness is significantly constrained in low-expert-data regimes. To mitigate this issue, previous works introduce ``*…

Cited by 0SourceScholar
2026

TraCeS: Learning Per-Timestep Constraint-Violation Credit from Sparse Trajectory-Level Labels

ICML 2026poster

Ensuring safe behavior in reinforcement learning (RL) is challenging when safety constraints are implicit and cannot be densely measured. In many settings, supervision is limited to coarse approvals or rejections of whole trajectories (e.g., whether a rollout remained within an unknown safety thresh…

Cited by 0SourceScholar
2025

IOSTOM: Offline Imitation Learning from Observations via State Transition Occupancy Matching

NeurIPS 2025poster

Offline Learning from Observations (LfO) focuses on enabling agents to imitate expert behavior using datasets that contain only expert state trajectories and separate transition data with suboptimal actions. This setting is both practical and critical in real-world scenarios where direct environment…

Cited by 0SourceScholar
2025

Leveraging Constraint Violation Signals for Action Constrained Reinforcement Learning

AAAI 2025technical

In many RL applications, ensuring an agent's actions adhere to constraints is crucial for safety. Most previous methods in Action-Constrained Reinforcement Learning (ACRL) employ a projection layer after the policy network to correct the action. However projection-based methods suffer from issues li…

2025

Offline Safe Reinforcement Learning Using Trajectory Classification

AAAI 2025technical

Offline safe reinforcement learning (RL) has emerged as a promising approach for learning safe behaviors without engaging in risky online interactions with the environment. Most existing methods in offline safe RL rely on cost constraints at each time step (derived from global cost constraints) and…

2024

Unified Training of Universal Time Series Forecasting Transformers

ICML 2024oral

Deep learning for time series forecasting has traditionally operated within a one-model-per-dataset framework, limiting its potential to leverage the game-changing impact of large pre-trained models. The concept of *universal forecasting*, emerging from pre-training on a vast collection of time seri…

2023

FlowPG: Action-constrained Policy Gradient with Normalizing Flows

NeurIPS 2023poster

Action-constrained reinforcement learning (ACRL) is a popular approach for solving safety-critical and resource-allocation related decision making problems. A major challenge in ACRL is to ensure agent taking a valid action satisfying constraints in each RL step. Commonly used approach of using a pr…

2023

Learning Deep Time-index Models for Time Series Forecasting

ICML 2023poster

Deep learning has been actively applied to time series forecasting, leading to a deluge of new methods, belonging to the class of historical-value models. Yet, despite the attractive properties of time-index models, such as being able to model the continuous nature of underlying time series dynamics…

2023

Planning and Learning for Non-markovian Negative Side Effects Using Finite State Controllers

AAAI 2023technical

Autonomous systems are often deployed in the open world where it is hard to obtain complete specifications of objectives and constraints. Operating based on an incomplete model can produce negative side effects (NSEs), which affect the safety and reliability of the system. We focus on mitigating NSE…

2023

Scalable and Globally Optimal Generalized L₁ K-center Clustering via Constraint Generation in Mixed Integer Linear Programming

AAAI 2023technical

The k-center clustering algorithm, introduced over 35 years ago, is known to be robust to class imbalance prevalent in many clustering problems and has various applications such as data summarization, document clustering, and facility location determination. Unfortunately, existing k-center algorith…

2022

CoST: Contrastive Learning of Disentangled Seasonal-Trend Representations for Time Series Forecasting

ICLR 2022poster

Deep learning has been actively studied for time series forecasting, and the mainstream paradigm is based on the end-to-end training of neural network architectures, ranging from classical LSTM/RNNs to more recent TCNs and Transformers. Motivated by the recent success of representation learning in c…

2022

Sample-Efficient Iterative Lower Bound Optimization of Deep Reactive Policies for Planning in Continuous MDPs

AAAI 2022technical

Recent advances in deep learning have enabled optimization of deep reactive policies (DRPs) for continuous MDP planning by encoding a parametric policy as a deep neural network and exploiting automatic differentiation in an end-to-end model-based gradient descent framework. This approach has proven…

2022

Using Constraint Programming and Graph Representation Learning for Generating Interpretable Cloud Security Policies

IJCAI 2022poster

Modern software systems rely on mining insights from business sensitive data stored in public clouds. A data breach usually incurs significant (monetary) loss for a commercial organization. Conceptually, cloud security heavily relies on Identity Access Management (IAM) policies that IT admins need t…

2018

Credit Assignment For Collective Multiagent RL With Global Rewards

NeurIPS 2018poster

Scaling decision theoretic planning to large multiagent systems is challenging due to uncertainty and partial observability in the environment. We focus on a multiagent planning model subclass, relevant to urban settings, where agent interactions are dependent on their ``collective influence'' on ea…

Cited by 133SourcePDFScholar
2017

Policy Gradient With Value Function Approximation For Collective Multiagent Planning

NeurIPS 2017poster

Decentralized (PO)MDPs provide an expressive framework for sequential decision making in a multiagent system. Given their computational complexity, recent research has focused on tractable yet practical subclasses of Dec-POMDPs. We address such a subclass called CDec-POMDP where the collective behav…

Cited by 73SourcePDFScholar
2016

Approximate Inference Using DC Programming For Collective Graphical Models

AISTATS 2016poster

Collective graphical models (CGMs) provide a framework for reasoning about a population of independent and identically distributed individuals when only noisy and aggregate observations are given. Previous approaches for inference in CGMs work on a junction-tree representation, thereby highly limit…

Cited by 18SourcePDFScholar