← Search

Krishna Balasubramanian

24 accepted papers

2026

Dependence-Aware Label Aggregation for LLM-as-a-Judge via Ising Models

ICML 2026poster

Large-scale AI evaluation increasingly relies on aggregating binary judgments from $K$ annotators, including LLMs used as judges. Most classical methods, e.g., Dawid-Skene or (weighted) majority voting, assume annotators are conditionally independent given the true label $Y\in\\{0,1\\}$, an assumpti…

Cited by 0SourceScholar
2026

Differentially Private Two-Stage Gradient Descent for Instrumental Variable Regression

ICLR 2026poster

We study instrumental variable regression (IVaR) under differential privacy constraints. Classical IVaR methods (like two-stage least squares regression) rely on solving moment equations that directly use sensitive covariates and instruments, creating significant risks of privacy leakage and posing…

Cited by 0SourceScholar
2026

Meta-Learning with Generalized Ridge Regression: High-dimensional Asymptotics, Optimality and Hyper-covariance Estimation

ICML 2026poster

Meta-learning involves training models on a variety of training tasks in a way that enables them to generalize well on new, unseen test tasks. In this work, we consider meta-learning within the framework of high-dimensional multivariate random-effects linear models and study generalized ridge-regres…

Cited by 0SourcecodeScholar
2026

Mirror Flow Matching with Heavy-Tailed Priors for Generative Modeling on Convex Domains

ICLR 2026poster

We study generative modeling on convex domains using flow matching and mirror maps, and identify two fundamental challenges. First, standard log-barrier mirror maps induce heavy-tailed dual distributions, leading to ill-posed dynamics. Second, coupling with Gaussian priors performs poorly when match…

Cited by 0SourceScholar
2025

Dense Associative Memory with Epanechnikov Energy

NeurIPS 2025spotlight

We propose a novel energy function for Dense Associative Memory (DenseAM) networks, the log-sum-ReLU (LSR), inspired by optimal kernel density estimation. Unlike the common log-sum-exponential (LSE) function, LSR is based on the Epanechnikov kernel and enables exact memory retrieval with exponential…

Cited by 0SourceScholar
2025

Improved Finite-Particle Convergence Rates for Stein Variational Gradient Descent

ICLR 2025oral

We provide finite-particle convergence rates for the Stein Variational Gradient Descent (SVGD) algorithm in the Kernelized Stein Discrepancy ($\KSD$) and Wasserstein-2 metrics. Our key insight is that the time derivative of the relative entropy between the joint density of $N$ particle locations and…

Cited by 3SourcePDFScholar
2025

Minimax Optimal Nonsmooth Nonparametric Regression via Fractional Laplacian Eigenmaps

UAI 2025

We develop minimax optimal estimators for nonparametric regression methods when the true regression function lies in an $L_2$-fractional Sobolev space with order $s\in (0,1)$. This function class is a Hilbert space lying between the space of square-integrable functions and the first-order Sobolev sp

Cited by 0SourcePDFScholar
2025

Restricted Spectral Gap Decomposition for Simulated Tempering Targeting Mixture Distributions

NeurIPS 2025poster

Simulated tempering is a widely used strategy for sampling from multimodal distributions. In this paper, we consider simulated tempering combined with an arbitrary local Markov chain Monte Carlo sampler and present a new decomposition theorem that provides a lower bound on the restricted spectral ga…

Cited by 0SourceScholar
2025

Transformers Handle Endogeneity in In-Context Linear Regression

ICLR 2025poster

We explore the capability of transformers to address endogeneity in in-context linear regression. Our main finding is that transformers inherently possess a mechanism to handle endogeneity effectively using instrumental variables (IV). First, we demonstrate that the transformer architecture can emul…

Cited by 2SourcePDFScholar
2024

A Separation in Heavy-Tailed Sampling: Gaussian vs. Stable Oracles for Proximal Samplers

NeurIPS 2024poster

We study the complexity of heavy-tailed sampling and present a separation result in terms of obtaining high-accuracy versus low-accuracy guarantees i.e., samplers that require only $\mathcal{O}(\log(1/\varepsilon))$ versus $\Omega(\text{poly}(1/\varepsilon))$ iterations to output a sample which is $…

Cited by 2SourcePDFScholar
2024

Adaptive and non-adaptive minimax rates for weighted Laplacian-Eigenmap based nonparametric regression

AISTATS 2024poster

We show both adaptive and non-adaptive minimax rates of convergence for a family of weighted Laplacian-Eigenmap based nonparametric regression methods, when the true regression function belongs to a Sobolev space and the sampling density is bounded from above and below. The adaptation methodology is…

Cited by 0SourcePDFScholar
2024

Stochastic Optimization Algorithms for Instrumental Variable Regression with Streaming Data

NeurIPS 2024poster

We develop and analyze algorithms for instrumental variable regression by viewing the problem as a conditional stochastic optimization problem. In the context of least-squares instrumental variable regression, our algorithms neither require matrix inversions nor mini-batches thereby providing a full…

Cited by 3SourcePDFScholar
2023

Decentralized Stochastic Bilevel Optimization with Improved per-Iteration Complexity

ICML 2023poster

Bilevel optimization recently has received tremendous attention due to its great success in solving important machine learning problems like meta learning, reinforcement learning, and hyperparameter optimization. Extending single-agent training on bilevel problems to the decentralized setting is a n…

Cited by 35SourcePDFScholar
2023

Forward-Backward Gaussian Variational Inference via JKO in the Bures-Wasserstein Space

ICML 2023poster

Variational inference (VI) seeks to approximate a target distribution $\pi$ by an element of a tractable family of distributions. Of key interest in statistics and machine learning is Gaussian VI, which approximates $\pi$ by minimizing the Kullback-Leibler (KL) divergence to $\pi$ over the space of…

Cited by 43SourcePDFScholar
2023

Towards Understanding the Dynamics of Gaussian-Stein Variational Gradient Descent

NeurIPS 2023poster

Stein Variational Gradient Descent (SVGD) is a nonparametric particle-based deterministic sampling algorithm. Despite its wide usage, understanding the theoretical properties of SVGD has remained a challenging problem. For sampling from a Gaussian target, the SVGD dynamics with a bilinear kernel wil…

Cited by 16SourcePDFScholar
2023

Winding Through: Crowd Navigation via Topological Invariance

RA-L 2023

We focus on robot navigation in crowded environments. The challenge of predicting the motion of a crowd around a robot makes it hard to ensure human safety and comfort. Recent approaches often employ end-to-end techniques for robot control or deep architectures for high-fidelity human motion predict

Cited by 41SourceScholar
2022

A Projection-free Algorithm for Constrained Stochastic Multi-level Composition Optimization

NeurIPS 2022accept

We propose a projection-free conditional gradient-type algorithm for smooth stochastic multi-level composition optimization, where the objective function is a nested composition of $T$ functions and the constraint set is a closed convex set. Our algorithm assumes access to noisy evaluations of the f…

Cited by 8SourcePDFScholar
2022

Constrained Stochastic Nonconvex Optimization with State-dependent Markov Data

NeurIPS 2022accept

We study stochastic optimization algorithms for constrained nonconvex stochastic optimization problems with Markovian data. In particular, we focus on the case when the transition kernel of the Markov chain is state-dependent. Such stochastic optimization problems arise in various machine learning p…

Cited by 11SourcePDFScholar
2022

High-probability bounds for robust stochastic Frank-Wolfe algorithm

UAI 2022poster

We develop and analyze robust Stochastic Frank-Wolfe type algorithms for projection-free stochastic convex optimization problems with heavy-tailed stochastic gradients. Existing works on the oracle complexity of such algorithms require a uniformly bounded variance assumption, and hold only in expect…

Cited by 3SourcePDFScholar
2021

An Analysis of Constant Step Size SGD in the Non-convex Regime: Asymptotic Normality and Bias

NeurIPS 2021poster

Structured non-convex learning problems, for which critical points have favorable statistical properties, arise frequently in statistical machine learning. Algorithmic convergence and statistical estimation rates are well-understood for such problems. However, quantifying the uncertainty associated…

Cited by 41SourcePDFScholar
2021

On Empirical Risk Minimization with Dependent and Heavy-Tailed Data

NeurIPS 2021poster

In this work, we establish risk bounds for Empirical Risk Minimization (ERM) with both dependent and heavy-tailed data-generating processes. We do so by extending the seminal works~\cite{pmlr-v35-mendelson14, mendelson2018learning} on the analysis of ERM with heavy-tailed but independent and identic…

Cited by 22SourcePDFScholar
2020

Fractal Gaussian Networks: A sparse random graph model based on Gaussian Multiplicative Chaos

ICML 2020poster

We propose a novel stochastic network model, called Fractal Gaussian Network (FGN), that embodies well-defined and analytically tractable fractal structures. Such fractal structures have been empirically observed in diverse applications. FGNs interpolate continuously between the popular purely rando…

Cited by 3SourcePDFScholar