← Search

Arunesh Sinha

15 accepted papers

2025

Bootstrapping Language Models with DPO Implicit Rewards

ICLR 2025poster

Human alignment in large language models (LLMs) is an active area of research. A recent groundbreaking work, direct preference optimization (DPO), has greatly simplified the process from past work in reinforcement learning from human feedback (RLHF) by bypassing the reward learning stage in RLHF. DP…

2025

On Minimizing Adversarial Counterfactual Error in Adversarial Reinforcement Learning

ICLR 2025poster

Deep Reinforcement Learning (DRL) policies are highly susceptible to adversarial noise in observations, which poses significant risks in safety-critical scenarios. The challenge inherent to adversarial perturbations is that by altering the information observed by the agent, the state becomes only pa…

2025

Semantic Loss Guided Data Efficient Supervised Fine Tuning for Safe Responses in LLMs

ICLR 2025poster

Large Language Models (LLMs) generating unsafe responses to toxic prompts is a significant issue in their applications. While various efforts aim to address this safety concern, previous approaches often demand substantial human data collection or rely on the less dependable option of using another…

Cited by 0SourcePDFScholar
2024

Handling Long and Richly Constrained Tasks through Constrained Hierarchical Reinforcement Learning

AAAI 2024technical

Safety in goal directed Reinforcement Learning (RL) settings has typically been handled through constraints over trajectories and have demonstrated good performance in primarily short horizon tasks. In this paper, we are specifically interested in the problem of solving temporally extended decision…

Cited by 0SourcePDFScholar
2024

Tackling Stackelberg Network Interdiction against a Boundedly Rational Adversary

IJCAI 2024poster

This work studies Stackelberg network interdiction games --- an important class of games in which a defender first allocates (randomized) defense resources to a set of critical nodes on a graph while an adversary chooses its path to attack these nodes accordingly. We consider a boundedly rational a…

Cited by 0SourcePDFScholar
2023

Behavioral Learning in Security Games: Threat of Multi-Step Manipulative Attacks

AAAI 2023technical

This paper studies the problem of multi-step manipulative attacks in Stackelberg security games, in which a clever attacker attempts to orchestrate its attacks over multiple time steps to mislead the defender's learning of the attacker's behavior. This attack manipulation eventually influences the d…

Cited by 0SourcePDFScholar
2023

Beyond NaN: Resiliency of Optimization Layers in the Face of Infeasibility

AAAI 2023technical

Prior work has successfully incorporated optimization layers as the last layer in neural networks for various problems, thereby allowing joint learning and planning in one neural network forward pass. In this work, we identify a weakness in such a set-up where inputs to the optimization layer lead t…

2023

Building a Personalized Messaging System for Health Intervention in Underprivileged Regions Using Reinforcement Learning

IJCAI 2023poster

This work builds an effective AI-based message generation system for diabetes prevention in rural areas, where the diabetes rate has been increasing at an alarming rate. The messages contain information about diabetes causes and complications and the impact of nutrition and fitness on preventing di…

Cited by 3SourcePDFScholar
2023

Generative Modelling of Stochastic Actions with Arbitrary Constraints in Reinforcement Learning

NeurIPS 2023poster

Many problems in Reinforcement Learning (RL) seek an optimal policy with large discrete multidimensional yet unordered action spaces; these include problems in randomized allocation of resources such as placements of multiple security resources and emergency response units, etc. A challenge in this…

2023

Securing Lifelines: Safe Delivery of Critical Services in Areas with Volatile Security Situation via a Stackelberg Game Approach

AAAI 2023technical

Vaccine delivery in under-resourced locations with security risks is not just challenging but also life threatening. The COVID pandemic and the need to vaccinate added even more urgency to this issue. Motivated by this problem, we propose a general framework to set-up limited temporary (vaccination)…

Cited by 1SourcePDFScholar
2022

Choices Are Not Independent: Stackelberg Security Games with Nested Quantal Response Models

AAAI 2022technical

The quantal response (QR) model is widely used in Stackelberg security games (SSG) to model a bounded rational adversary. The QR model is a model of human response from among a large variety of prominent models known as discrete choice models. QR is the simplest type of discrete choice models and do…

Cited by 2SourcePDFScholar
2022

Multiscale Generative Models: Improving Performance of a Generative Model Using Feedback from Other Dependent Generative Models

AAAI 2022technical

Realistic fine-grained multi-agent simulation of real-world complex systems is crucial for many downstream tasks such as reinforcement learning. Recent work has used generative models (GANs in particular) for providing high-fidelity simulation of real-world systems. However, such generative models a…

2022

Scalable Distributional Robustness in a Class of Non-Convex Optimization with Guarantees

NeurIPS 2022accept

Distributionally robust optimization (DRO) has shown a lot of promise in providing robustness in learning as well as sample-based optimization problems. We endeavor to provide DRO solutions for a class of sum of fractionals, non-convex optimization which is used for decision making in prominent area…

Cited by 2SourcePDFScholar