← Search

Soummya Kar

18 accepted papers

2026

Decentralized Nonconvex Optimization under Heavy-Tailed Noise: Normalization and Optimal Convergence

ICLR 2026poster

Heavy-tailed noise in nonconvex stochastic optimization has garnered increasing research interest, as empirical studies, including those on training attention models, suggest it is a more realistic gradient noise condition. This paper studies first-order nonconvex stochastic optimization under heavy…

Cited by 0SourceScholar
2026

Enhancing Complex Symbolic Logical Rea­soning of Large Language Models via Sparse Multi-Agent Debate

ICLR 2026poster

Large language models (LLMs) struggle with complex logical reasoning. Previous work has primarily explored single-agent methods, with their performance remains fundamentally limited by the capabilities of a single model. To our knowledge, this paper first introduce a multi-agent approach specificall…

Cited by 0SourcecodeScholar
2026

TEST-TIME SCALING IN DIFFUSION LLMS VIA HIDDEN SEMI-AUTOREGRESSIVE EXPERTS

ICLR 2026poster

Diffusion-based large language models (dLLMs) are trained to model extreme flexibility/dependence in the data-distribution; however, how to best utilize this at inference time remains an open problem. In this work, we uncover an interesting property of these models: dLLMs {trained on textual data} i…

Cited by 0SourceScholar
2025

High-probability Convergence Bounds for Online Nonlinear Stochastic Gradient Descent under Heavy-tailed Noise

AISTATS 2025poster

We study high-probability convergence in online learning, in the presence of heavy-tailed noise. To combat the heavy tails, a general framework of nonlinear SGD methods is considered, subsuming several popular nonlinearities like sign, quantization, component-wise and joint clipping. In our work the…

Cited by 0SourceScholar
2023

Large deviations rates for stochastic gradient descent with strongly convex functions

AISTATS 2023poster

Recent works have shown that high probability metrics with stochastic gradient descent (SGD) exhibit informativeness and in some cases advantage over the commonly adopted mean-square error-based ones. In this work we provide a formal framework for the study of general high probability bounds with SG…

Cited by 7SourcePDFScholar
2021

A Decentralized Variance-Reduced Method for Stochastic Optimization Over Directed Graphs

ICASSP 2021accepted

In this paper, we propose a decentralized first-order stochastic optimization method Push-SAGA for finite-sum minimization over a strongly connected directed graph. This method features local variance reduction to remove the uncertainty caused by random sampling of the local gradients, global gradie…

Cited by 4SourceScholar
2021

A Hybrid Variance-Reduced Method for Decentralized Stochastic Non-Convex Optimization

ICML 2021spotlight

This paper considers decentralized stochastic optimization over a network of $n$ nodes, where each node possesses a smooth non-convex local cost function and the goal of the networked nodes is to find an $\epsilon$-accurate first-order stationary point of the sum of the local costs. We focus on an o…

Cited by 46SourcePDFScholar
2019

Towards Gradient Free and Projection Free Stochastic Optimization

AISTATS 2019poster

This paper focuses on the problem of \emph{constrained} \emph{stochastic} optimization. A zeroth order Frank-Wolfe algorithm is proposed, which in addition to the projection-free nature of the vanilla Frank-Wolfe algorithm makes it gradient free. Under convexity and smoothness assumption, we show th…

Cited by 48SourcePDFScholar
2017

Convergence analysis of the information matrix in Gaussian Belief Propagation

ICASSP 2017accepted

Gaussian belief propagation (BP) has been widely used for distributed estimation in large-scale networks such as the smart grid, communication networks, and social networks, where local meansurements/observations are scattered over a wide geographical area. However, the convergence of Gaussian BP is…

Cited by 0SourceScholar
2017

Fast path localization on graphs via multiscale Viterbi decoding

ICASSP 2017accepted

We consider a problem of localizing the destination of an activated path signal supported on a graph. An “activated path signal” is a graph signal that evolves over time that can be viewed as the trajectory of a moving agent. We show that by combining dynamic programming and graph partitioning, the…

Cited by 0SourceScholar