← Search

Arvind Easwaran

3 accepted papers

2026

CAPO: A Unified Policy Gradient Approach for Reward and Cost Optimization in Safe Reinforcement Learning (Student Abstract)

AAAI 2026technical

In safe reinforcement learning (SRL), there exists an inherent conflict between maximizing reward and minimizing cost. We propose a novel approach that effectively resolve the conflict between maximizing reward and minimizing cost in joint optimization.When the cost exceeds the threshold, we perform

Cited by 0SourcePDFScholar
2025

Adaptive Multi-prompt Contrastive Network for Few-shot Out-of-distribution Detection

ICML 2025spotlight

Out-of-distribution (OOD) detection attempts to distinguish outlier samples to prevent models trained on the in-distribution (ID) dataset from producing unavailable outputs. Most OOD detection methods require many ID samples for training, which seriously limits their real-world applications. To this…

Cited by 0SourcePDFScholar
2025

Guaranteeing Out-Of-Distribution Detection in Deep RL via Transition Estimation

AAAI 2025technical

An issue concerning the use of deep reinforcement learning (RL) agents is whether they can be trusted to perform reliably when deployed, as training environments may not reflect real-life environments. Anticipating instances outside their training scope, learning-enabled systems are often equipped w…

Cited by 0SourcePDFScholar