← Search

Panganamala Kumar

9 accepted papers

2026

Diffusion Blend: Inference-Time Multi-Preference Alignment for Diffusion Models

ICLR 2026poster

Reinforcement learning (RL) algorithms have been used recently to align diffusion models with downstream objectives such as aesthetic quality and text-image consistency by fine-tuning them to maximize a single reward function under a fixed KL regularization. However, this approach is inherently rest…

Cited by 0SourcecodeScholar
2024

Finite Time Logarithmic Regret Bounds for Self-Tuning Regulation

ICML 2024poster

We establish the first finite-time logarithmic regret bounds for the self-tuning regulation problem. We introduce a modified version of the certainty equivalence algorithm, which we call PIECE, that clips inputs in addition to utilizing probing inputs for exploration. We show that it has a $C \log T…

Cited by 0SourcePDFScholar
2024

Is O(log N) practical? Near-Equivalence Between Delay Robustness and Bounded Regret in Bandits and RL

NeurIPS 2024poster

Interactive decision making, encompassing bandits, contextual bandits, and reinforcement learning, has recently been of interest to theoretical studies of experimentation design and recommender system algorithm research. One recent finding in this area is that the well-known Graves-Lai constant bein…

Cited by 0SourcePDFScholar
2023

Natural Actor-Critic for Robust Reinforcement Learning with Function Approximation

NeurIPS 2023poster

We study robust reinforcement learning (RL) with the goal of determining a well-performing policy that is robust against model mismatch between the training simulator and the testing environment. Previous policy-based robust RL algorithms mainly focus on the tabular setting under uncertainty sets th…

2023

Provably Fast Convergence of Independent Natural Policy Gradient for Markov Potential Games

NeurIPS 2023poster

This work studies an independent natural policy gradient (NPG) algorithm for the multi-agent reinforcement learning problem in Markov potential games. It is shown that, under mild technical assumptions and the introduction of the \textit{suboptimality gap}, the independent NPG method with an oracle…

2022

Anchor-Changing Regularized Natural Policy Gradient for Multi-Objective Reinforcement Learning

NeurIPS 2022accept

We study policy optimization for Markov decision processes (MDPs) with multiple reward value functions, which are to be jointly optimized according to given criteria such as proportional fairness (smooth concave scalarization), hard constraints (constrained MDP), and max-min trade-off. We propose an…

2022

Augmented RBMLE-UCB Approach for Adaptive Control of Linear Quadratic Systems

NeurIPS 2022accept

We consider the problem of controlling an unknown stochastic linear system with quadratic costs -- called the adaptive LQ control problem. We re-examine an approach called ``Reward-Biased Maximum Likelihood Estimate'' (RBMLE) that was proposed more than forty years ago, and which predates the ``Uppe…

Cited by 13SourcePDFScholar
2022

Learning from Few Samples: Transformation-Invariant SVMs with Composition and Locality at Multiple Scales

NeurIPS 2022accept

Motivated by the problem of learning with small sample sizes, this paper shows how to incorporate into support-vector machines (SVMs) those properties that have made convolutional neural networks (CNNs) successful. Particularly important is the ability to incorporate domain knowledge of invariances,…

2021

Learning Policies with Zero or Bounded Constraint Violation for Constrained MDPs

NeurIPS 2021poster

We address the issue of safety in reinforcement learning. We pose the problem in an episodic framework of a constrained Markov decision process. Existing results have shown that it is possible to achieve a reward regret of $\tilde{\mathcal{O}}(\sqrt{K})$ while allowing an $\tilde{\mathcal{O}}(\sqrt{…

Cited by 95SourcePDFScholar