← Search

Anit Kumar Sahu

10 accepted papers

2026

Direct Preference Optimization for Primitive-Enabled Hierarchical RL: A Bilevel Approach

ICLR 2026poster

Hierarchical reinforcement learning (HRL) enables agents to solve complex, long-horizon tasks by decomposing them into manageable sub-tasks. However, HRL methods face two fundamental challenges: (i) non-stationarity caused by the evolving lower-level policy during training, which destabilizes higher…

Cited by 0SourceScholar
2026

TEST-TIME SCALING IN DIFFUSION LLMS VIA HIDDEN SEMI-AUTOREGRESSIVE EXPERTS

ICLR 2026poster

Diffusion-based large language models (dLLMs) are trained to model extreme flexibility/dependence in the data-distribution; however, how to best utilize this at inference time remains an open problem. In this work, we uncover an interesting property of these models: dLLMs {trained on textual data} i…

Cited by 0SourceScholar
2024

Get more for less: Principled Data Selection for Warming Up Fine-Tuning in LLMs

ICLR 2024poster

This work focuses on leveraging and selecting from vast, unlabeled, open data to *pre-fine-tune* a pre-trained language model. The goal is to minimize the need for costly domain-specific data for subsequent fine-tuning while achieving desired performance levels. While many data selection algorithms…

Cited by 15SourcePDFScholar
2023

Federated Self-Learning with Weak Supervision for Speech Recognition

ICASSP 2023accepted

Automatic speech recognition (ASR) models with low-footprint are increasingly being deployed on edge devices for conversational agents, which enhances privacy. We study the problem of federated continual incremental learning for recurrent neural network-transducer (RNN-T) ASR models in the privacy-e…

Cited by 0SourceScholar
2023

Performance Scaling via Optimal Transport: Enabling Data Selection from Partially Revealed Sources

NeurIPS 2023poster

Traditionally, data selection has been studied in settings where all samples from prospective sources are fully revealed to a machine learning developer. However, in practical data exchange scenarios, data providers often reveal only a limited subset of samples before an acquisition decision is made…

Cited by 14SourcePDFScholar
2022

Federated Learning Challenges and Opportunities: An Outlook

ICASSP 2022accepted

Federated learning (FL) has been developed as a promising framework to leverage the resources of edge devices, enhance customers’ privacy, comply with regulations, and reduce development costs. Although many methods and applications have been developed for FL, several critical challenges for practic…

Cited by 0SourceScholar
2022

Self-Aware Personalized Federated Learning

NeurIPS 2022accept

In the context of personalized federated learning (FL), the critical challenge is to balance local model improvement and global model tuning when the personal and global objectives may not be exactly aligned. Inspired by Bayesian hierarchical models, we develop a self-aware personalized FL method wh…

Cited by 28SourcePDFScholar
2019

Towards Gradient Free and Projection Free Stochastic Optimization

AISTATS 2019poster

This paper focuses on the problem of \emph{constrained} \emph{stochastic} optimization. A zeroth order Frank-Wolfe algorithm is proposed, which in addition to the projection-free nature of the vanilla Frank-Wolfe algorithm makes it gradient free. Under convexity and smoothness assumption, we show th…

Cited by 48SourcePDFScholar