← Search

Prashant Khanduri

17 accepted papers

2026

Not All Tokens Are Meant to Be Forgotten

AAAI 2026technical

Large Language Models (LLMs), pre-trained on massive text corpora, exhibit remarkable human-level language understanding, reasoning, and decision-making abilities. However, they tend to memorize unwanted information, such as private or copyrighted content, raising significant privacy and legal conce

Cited by 0SourcePDFScholar
2026

WalkGPT: Grounded Vision-Language Conversation with Depth-Aware Segmentation for Pedestrian Navigation

CVPR 2026

Ensuring accessible pedestrian navigation requires reasoning about both semantic and spatial aspects of complex urban scenes, a challenge that existing Large Vision-Language Models (LVLMs) struggle to meet. Although these models can describe visual content, their lack of explicit grounding leads to

Cited by 0SourcecodeScholar
2024

Fairness-aware Vision Transformer via Debiased Self-Attention

ECCV 2024poster

"Vision Transformer (ViT) has recently gained significant attention in solving computer vision (CV) problems due to its capability of extracting informative features and modeling long-range dependencies through the attention mechanism. Whereas recent works have explored the trustworthiness of ViT, i…

2024

Understanding Server-Assisted Federated Learning in the Presence of Incomplete Client Participation

ICML 2024poster

Existing works in federated learning (FL) often assume either full client or uniformly distributed client participation. However, in reality, some clients may never participate in FL training (aka incomplete client participation) due to various system heterogeneity factors. A popular solution is the…

Cited by 1SourcePDFScholar
2023

An Implicit Gradient Method for Constrained Bilevel Problems Using Barrier Approximation

ICASSP 2023accepted

In this work, we propose algorithms for solving a class of Bilevel Optimization (BLO) problems, with applications in areas such as signal processing, networking and machine learning. Specifically, we develop a novel barrier-based gradient approximation algorithm that transforms the constrained BLO p…

Cited by 0SourceScholar
2023

FedAvg Converges to Zero Training Loss Linearly for Overparameterized Multi-Layer Neural Networks

ICML 2023poster

Federated Learning (FL) is a distributed learning paradigm that allows multiple clients to learn a joint model by utilizing privately held data at each client. Significant research efforts have been devoted to develop advanced algorithms that deal with the situation where the data at individual clie…

Cited by 8SourcePDFScholar
2023

Linearly Constrained Bilevel Optimization: A Smoothed Implicit Gradient Approach

ICML 2023poster

This work develops analysis and algorithms for solving a class of bilevel optimization problems where the lower-level (LL) problems have linear constraints. Most of the existing approaches for constrained bilevel problems rely on value function-based approximate reformulations, which suffer from iss…

Cited by 22SourcePDFScholar
2023

Prometheus: Taming Sample and Communication Complexities in Constrained Decentralized Stochastic Bilevel Learning

ICML 2023poster

In recent years, decentralized bilevel optimization has gained significant attention thanks to its versatility in modeling a wide range of multi-agent learning problems, such as multi-agent reinforcement learning and multi-agent meta-learning. However, one unexplored and fundamental problem in this…

Cited by 6SourcePDFScholar
2022

An Implicit Gradient-Type Method for Linearly Constrained Bilevel Problems

ICASSP 2022accepted

In this work, we develop an implicit gradient-type (IG-AL) algorithm for bilevel optimization with strongly convex linear inequality constrained lower-level problems. Many learning problems of interest, including problems in distributed optimization, machine learning, economics, and transport resear…

Cited by 0SourceScholar
2022

Decentralized Learning for Overparameterized Problems: A Multi-Agent Kernel Approximation Approach

ICLR 2022poster

This work develops a novel framework for communication-efficient distributed learning where the models to be learned are overparameterized. We focus on a class of kernel learning problems (which includes the popular neural tangent kernel (NTK) learning as a special case) and propose a novel {\it mul…

Cited by 0SourcePDFScholar
2022

Revisiting and Advancing Fast Adversarial Training Through The Lens of Bi-Level Optimization

ICML 2022spotlight

Adversarial training (AT) is a widely recognized defense mechanism to gain the robustness of deep neural networks against adversarial attacks. It is built on min-max optimization (MMO), where the minimizer (i.e., defender) seeks a robust model to minimize the worst-case training loss in the presence…

2021

A Near-Optimal Algorithm for Stochastic Bilevel Optimization via Double-Momentum

NeurIPS 2021poster

This paper proposes a new algorithm -- the \underline{S}ingle-timescale Do\underline{u}ble-momentum \underline{St}ochastic \underline{A}pprox\underline{i}matio\underline{n} (SUSTAIN) -- for tackling stochastic unconstrained bilevel optimization problems. We focus on bilevel problems where the lower…

Cited by 147SourcePDFScholar
2021

STEM: A Stochastic Two-Sided Momentum Algorithm Achieving Near-Optimal Sample and Communication Complexities for Federated Learning

NeurIPS 2021poster

Federated Learning (FL) refers to the paradigm where multiple worker nodes (WNs) build a joint model by using local data. Despite extensive research, for a generic non-convex FL problem, it is not clear, how to choose the WNs' and the server's update directions, the minibatch sizes, and the local up…

Cited by 73SourcePDFScholar
2020

On Distributed Stochastic Gradient Descent for Nonconvex Functions in the Presence of Byzantines

ICASSP 2020accepted

We consider the distributed stochastic optimization problem of minimizing a nonconvex function f in an adversarial setting. All the w worker nodes in the network are expected to send their stochastic gradient vectors to the fusion center (or server). However, some (at most α-fraction) of the nodes m…

Cited by 0SourceScholar
2018

On Sequential Random Distortion Testing of Non-Stationary Processes

ICASSP 2018accepted

Random distortion testing (RDT) addresses the problem of testing whether or not a random signal, Ξ, deviates by more than a specified tolerance, τ, from a fixed value, ξ <sub xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">0</sub> [1]. The test is nonparamet…

Cited by 1SourceScholar