← Search

Parameswaran Kamalaruban

15 accepted papers

2026

DMAP: A Distribution Map for Text

ICLR 2026poster

Large Language Models (LLMs) are a powerful tool for statistical text analysis, with derived sequences of next-token probability distributions offering a wealth of information. Extracting this signal typically relies on metrics such as perplexity, which do not adequately account for context; how one…

Cited by 0SourceScholar
2025

Corruption Robust Offline Reinforcement Learning with Human Feedback

AISTATS 2025oral

We study data corruption robustness for reinforcement learning with human feedback (RLHF) in an offline setting. Given an offline dataset of pairs of trajectories along with feedback about human preferences, an $\varepsilon$-fraction of the pairs is corrupted (e.g., feedback flipped or trajectory f…

Cited by 0SourceScholar
2025

Curriculum Design for Trajectory-Constrained Agent: Compressing Chain-of-Thought Tokens in LLMs

NeurIPS 2025poster

Training agents to operate under strict constraints during deployment, such as limited resource budgets or stringent safety requirements, presents significant challenges, especially when these constraints render the task complex. In this work, we propose a curriculum learning strategy that gradually…

Cited by 0SourceScholar
2025

Inference-Time Personalized Alignment with a Few User Preference Queries

NeurIPS 2025poster

We study the problem of aligning a generative model's response with a user's preferences. Recent works have proposed several different formulations for personalized alignment; however, they either require a large amount of user preference queries or require that the preference be explicitly specifie…

Cited by 0SourceScholar
2025

Learning Personalized Decision Support Policies

AAAI 2025technical

Individual human decision-makers may benefit from different forms of support to improve decision outcomes, but when will each form of support yield better outcomes? In this work, we posit that personalizing access to decision support tools can be an effective mechanism for instantiating the appropri…

Cited by 15SourcePDFScholar
2025

Policy Teaching via Data Poisoning in Learning from Human Preferences

AISTATS 2025poster

We study data poisoning attacks in learning from human preferences. More specifically, we consider the problem of teaching/enforcing a target policy $\pi^\dagger$ by synthesizing preference data. We seek to understand the susceptibility of different preference-based learning paradigms to poisoned pr…

Cited by 0SourceScholar
2024

Adversarially Robust Decision Transformer

NeurIPS 2024poster

Decision Transformer (DT), as one of the representative Reinforcement Learning via Supervised Learning (RvS) methods, has achieved strong performance in offline learning tasks by leveraging the powerful Transformer architecture for sequential decision-making. However, in adversarial environments, th…

2024

Proximal Curriculum with Task Correlations for Deep Reinforcement Learning

IJCAI 2024poster

Curriculum design for reinforcement learning (RL) can speed up an agent's learning process and help it learn to perform well on complex tasks. However, existing techniques typically require domain-specific hyperparameter tuning, involve expensive optimization procedures for task selection, or are su…

2024

Reward Model Learning vs. Direct Policy Optimization: A Comparative Analysis of Learning from Human Preferences

ICML 2024poster

In this paper, we take a step towards a deeper understanding of learning from human preferences by systematically comparing the paradigm of reinforcement learning from human feedback (RLHF) with the recently proposed paradigm of direct preference optimization (DPO). We focus our attention on the cla…

Cited by 10SourcePDFScholar
2022

Exploration-Guided Reward Shaping for Reinforcement Learning under Sparse Rewards

NeurIPS 2022accept

We study the problem of reward shaping to accelerate the training process of a reinforcement learning agent. Existing works have considered a number of different reward shaping formulations; however, they either require external domain knowledge or fail in environments with extremely sparse rewards.…

Cited by 63SourcePDFScholar
2021

Curriculum Design for Teaching via Demonstrations: Theory and Applications

NeurIPS 2021poster

We consider the problem of teaching via demonstrations in sequential decision-making settings. In particular, we study how to design a personalized curriculum over demonstrations to speed up the learner's convergence. We provide a unified curriculum strategy for two popular learner models: Maximum C…

2021

Explicable Reward Design for Reinforcement Learning Agents

NeurIPS 2021poster

We study the design of explicable reward functions for a reinforcement learning agent while guaranteeing that an optimal policy induced by the function belongs to a set of target policies. By being explicable, we seek to capture two properties: (a) informativeness so that the rewards speed up the ag…

2021

Robust Inverse Reinforcement Learning under Transition Dynamics Mismatch

NeurIPS 2021poster

We study the inverse reinforcement learning (IRL) problem under a transition dynamics mismatch between the expert and the learner. Specifically, we consider the Maximum Causal Entropy (MCE) IRL learner model and provide a tight upper bound on the learner's performance degradation based on the $\ell_…

2020

Robust Reinforcement Learning via Adversarial training with Langevin Dynamics

NeurIPS 2020poster

We introduce a \emph{sampling} perspective to tackle the challenging task of training robust Reinforcement Learning (RL) agents. Leveraging the powerful Stochastic Gradient Langevin Dynamics, we present a novel, scalable two-player RL algorithm, which is a sampling variant of the two-player policy g…

Cited by 73SourcePDFScholar