← Search

Subhojyoti Mukherjee

13 accepted papers

2026

Stepwise Credit Assignment for GRPO on Flow-Matching Models

CVPR 2026

Flow-GRPO successfully applies reinforcement learning to flow models, but uses uniform credit assignment across all steps. This ignores the temporal structure of diffusion generation: early steps determine composition and content (low-frequency structure), while late steps resolve details and textur

Cited by 0SourceScholar
2025

From Selection to Generation: A Survey of LLM-based Active Learning

ACL 2025long

Active Learning (AL) has been a powerful paradigm for improving model efficiency and performance by selecting the most informative data points for labeling and training. In recent active learning frameworks, Large Language Models (LLMs) have been employed not only for selection but also for generati…

Cited by 0SourcePDFScholar
2025

Logits are All We Need to Adapt Closed Models

ICML 2025poster

Many commercial Large Language Models (LLMs) are often closed-source, limiting developers to prompt tuning for aligning content generation with specific applications. While these models currently do not provide access to token logits, we argue that if such access were available, it would enable more…

2025

Offline RL by Reward-Weighted Fine-Tuning for Conversation Optimization

NeurIPS 2025poster

Offline reinforcement learning (RL) is a variant of RL where the policy is learned from a previously collected dataset of trajectories and rewards. In our work, we propose a practical approach to offline RL with large language models (LLMs). We recast the problem as reward-weighted fine-tuning, whic…

Cited by 0SourceScholar
2024

Optimal Design for Human Preference Elicitation

NeurIPS 2024poster

Learning of preference models from human feedback has been central to recent advances in artificial intelligence. Motivated by the cost of obtaining high-quality human annotations, we study efficient human preference elicitation for learning preference models. The key idea in our work is to generali…

Cited by 5SourcePDFScholar
2024

SPEED: Experimental Design for Policy Evaluation in Linear Heteroscedastic Bandits

AISTATS 2024poster

In this paper, we study the problem of optimal data collection for policy evaluation in linear bandits. In policy evaluation, we are given a \textit{target} policy and asked to estimate the expected reward it will obtain when executed in a multi-armed bandit environment. Our work is the first work t…

Cited by 7SourcePDFScholar
2024

SaVeR: Optimal Data Collection Strategy for Safe Policy Evaluation in Tabular MDP

ICML 2024poster

In this paper, we study safe data collection for the purpose of policy evaluation in tabular Markov decision processes (MDPs). In policy evaluation, we are given a target policy and asked to estimate the expected cumulative reward it will obtain. Policy evaluation requires data and we are interested…

Cited by 2SourcePDFScholar
2023

Multi-task Representation Learning for Pure Exploration in Bilinear Bandits

NeurIPS 2023poster

We study multi-task representation learning for the problem of pure exploration in bilinear bandits. In bilinear bandits, an action takes the form of a pair of arms from two different entity types and the reward is a bilinear function of the known feature vectors of the arms. In the \textit{multi-ta…

Cited by 7SourcePDFScholar
2022

Chernoff Sampling for Active Testing and Extension to Active Regression

AISTATS 2022poster

Active learning can reduce the number of samples needed to perform a hypothesis test and to estimate the parameters of a model. In this paper, we revisit the work of Chernoff that described an asymptotically optimal algorithm for performing a hypothesis test. We obtain a novel sample complexity boun…

Cited by 14SourcePDFScholar
2022

Nearly Optimal Algorithms for Level Set Estimation

AISTATS 2022poster

The level set estimation problem seeks to find all points in a domain $\mathcal{X}$ where the value of an unknown function $f:\mathcal{X}\rightarrow \mathbb{R}$ exceeds a threshold $\alpha$. The estimation is based on noisy function evaluations that may be acquired at sequentially and adaptively cho…

Cited by 27SourcePDFScholar
2022

ReVar: Strengthening policy evaluation via reduced variance sampling

UAI 2022poster

This paper studies the problem of data collection for policy evaluation in Markov decision processes (MDPs). In policy evaluation, we are given a \textit{target} policy and asked to estimate the expected cumulative reward it will obtain in an environment formalized as an MDP. We develop theory for o…

Cited by 17SourcePDFScholar
2021

A Unified Approach to Translate Classical Bandit Algorithms to Structured Bandits

ICASSP 2021accepted

We consider a finite-armed structured bandit problem in which mean rewards of different arms are known functions of a common hidden parameter θ <sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">*</sup> . This problem setting subsumes several previously st…

Cited by 0SourceScholar