← Search

Dhruv Rohatgi

13 accepted papers

2026

Taming Imperfect Process Verifiers: A Sampling Perspective on Backtracking

ICLR 2026poster

Test-time algorithms that combine the *generative* power of language models with *process verifiers* that assess the quality of partial generations offer a promising lever for eliciting new reasoning capabilities, but the algorithmic design space and computational scaling properties of such approach…

Cited by 0SourceScholar
2025

Self-Improvement in Language Models: The Sharpening Mechanism

ICLR 2025oral

Recent work in language modeling has raised the possibility of “self-improvement,” where an LLM evaluates and refines its own generations to achieve higher performance without external feedback. It is impossible for this self-improvement to create information that is not already in the model, so why…

Cited by 5SourcePDFScholar
2025

To Distill or Decide? Understanding the Algorithmic Trade-off in Partially Observable RL

NeurIPS 2025spotlight

Partial observability is a notorious challenge in reinforcement learning (RL), due to the need to learn complex, history-dependent policies. Recent empirical successes have used *privileged expert distillation* -- which leverages availability of latent state information during training (e.g., from…

Cited by 0SourceScholar
2025

Towards characterizing the value of edge embeddings in Graph Neural Networks

ICML 2025poster

Graph neural networks (GNNs) are the dominant approach to solving machine learning problems defined over graphs. Despite much theoretical and empirical work in recent years, our understanding of finer-grained aspects of architectural design for GNNs remains impoverished. In this paper, we consider t…

Cited by 1SourcePDFScholar
2023

Provable benefits of score matching

NeurIPS 2023spotlight

Score matching is an alternative to maximum likelihood (ML) for estimating a probability distribution parametrized up to a constant of proportionality. By fitting the ''score'' of the distribution, it sidesteps the need to compute this constant of proportionality (which is often intractable). While…

Cited by 16SourcePDFScholar
2022

Learning in Observable POMDPs, without Computationally Intractable Oracles

NeurIPS 2022accept

Much of reinforcement learning theory is built on top of oracles that are computationally hard to implement. Specifically for learning near-optimal policies in Partially Observable Markov Decision Processes (POMDPs), existing algorithms either need to make strong assumptions about the model dynamics…

Cited by 44SourcePDFScholar
2022

Lower Bounds on Randomly Preconditioned Lasso via Robust Sparse Designs

NeurIPS 2022accept

Sparse linear regression with ill-conditioned Gaussian random covariates is widely believed to exhibit a statistical/computational gap, but there is surprisingly little formal evidence for this belief. Recent work has shown that, for certain covariance matrices, the broad class of Preconditioned Las…

Cited by 5SourcePDFScholar
2020

Constant-Expansion Suffices for Compressed Sensing with Generative Priors

NeurIPS 2020spotlight

Generative neural networks have been empirically found very promising in providing effective structural priors for compressed sensing, since they can be trained to span low-dimensional data manifolds in high-dimensional signal spaces. Despite the non-convexity of the resulting optimization problem,…

Cited by 20SourcePDFScholar