← Search

Abhishek Kumar

38 accepted papers

2025

Exponentially Weighted Instance-Aware Repeat Factor Sampling for Long-Tailed Object Detection Model Training in Unmanned Aerial Vehicles Surveillance Scenarios

IROS 2025

Object detection models often struggle with class imbalance, where rare categories appear significantly less frequently than common ones. Existing sampling-based rebalancing strategies, such as Repeat Factor Sampling (RFS) and Instance-Aware Repeat Factor Sampling (IRFS), mitigate this issue by adju

Cited by 0SourcecodeScholar
2025

FairGen: Controlling Sensitive Attributes for Fair Generations in Diffusion Models via Adaptive Latent Guidance

EMNLP 2025

Text-to-image diffusion models often exhibit biases toward specific demographic groups, such as generating more males than females when prompted to generate images of engineers, raising ethical concerns and limiting their adoption. In this paper, we tackle the challenge of mitigating generation bias

2025

Frequency Agnostic Tissue Characterization in Ultrasound Imaging using Backscattered Signal Statistics

ICASSP 2025accepted

Ultrasound (US) imaging-based tissue characterization (TC) is a vital tool for improving diagnostic accuracy by assessing tissue properties. However, existing methods often lack generalizability across varying US acquisition frequencies. This paper introduces a frequency-agnostic TC method that esti…

Cited by 0SourceScholar
2025

RB-Modulation: Training-Free Stylization using Reference-Based Modulation

ICLR 2025oral

We propose Reference-Based Modulation (RB-Modulation), a new plug-and-play solution for training-free personalization of diffusion models. Existing training-free approaches exhibit difficulties in (a) style extraction from reference images in the absence of additional style or content text descripti…

2025

Wanda++: Pruning Large Language Models via Regional Gradients

ACL 2025finding

Large Language Models (LLMs) pruning seeks to remove unimportant weights for inference speedup with minimal accuracy impact. However, existing methods often suffer from accuracy degradation without full-model sparsity-aware fine-tuning. This paper presents Wanda++, a novel pruning framework that out…

Cited by 0SourcePDFScholar
2024

Beyond First-Order Tweedie: Solving Inverse Problems using Latent Diffusion

CVPR 2024poster

Sampling from the posterior distribution in latent diffusion models for inverse problems is computationally challenging. Existing methods often rely on Tweedie's first-order moments that tend to induce biased results. Second-order approximations are computationally prohibitive making standard revers…

Cited by 26SourcePDFScholar
2024

Confidence Under the Hood: An Investigation into the Confidence-Probability Alignment in Large Language Models

ACL 2024long

As the use of Large Language Models (LLMs) becomes more widespread, understanding their self-evaluation of confidence in generated responses becomes increasingly important as it is integral to the reliability of the output of these models. We introduce the concept of Confidence-Probability Alignment…

2024

Found in the middle: Calibrating Positional Attention Bias Improves Long Context Utilization

ACL 2024findings

Large language models (LLMs), even when specifically trained to process long input contexts, struggle to capture relevant information located in the middle of their input. This phenomenon has been known as the lost-in-the-middle problem. In this work, we make three contributions. First, we set out t…

2024

Small-scale proxies for large-scale Transformer training instabilities

ICLR 2024oral

Teams that have trained large Transformer-based models have reported training instabilities at large scale that did not appear when training with the same hyperparameters at smaller scales. Although the causes of such instabilities are of scientific interest, the amount of resources required to repr…

Cited by 79SourcePDFScholar
2024

Subtle Biases Need Subtler Measures: Dual Metrics for Evaluating Representative and Affinity Bias in Large Language Models

ACL 2024long

Research on Large Language Models (LLMs) has often neglected subtle biases that, although less apparent, can significantly influence the models’ outputs toward particular social narratives. This study addresses two such biases within LLMs: representative bias, which denotes a tendency of LLMs to gen…

2024

The Group Robustness is in the Details: Revisiting Finetuning under Spurious Correlations

NeurIPS 2024poster

Modern machine learning models are prone to over-reliance on spurious correlations, which can often lead to poor performance on minority groups. In this paper, we identify surprising and nuanced behavior of finetuned models on worst-group accuracy via comprehensive experiments on four well-establish…

2023

Distributionally Robust Post-hoc Classifiers under Prior Shifts

ICLR 2023poster

The generalization ability of machine learning models degrades significantly when the test distribution shifts away from the training distribution. We investigate the problem of training models that are robust to shifts caused by changes in the distribution of class-priors or group-priors. The prese…

2023

Towards Last-layer Retraining for Group Robustness with Fewer Annotations

NeurIPS 2023poster

Empirical risk minimization (ERM) of neural networks is prone to over-reliance on spurious correlations and poor generalization on minority groups. The recent deep feature reweighting (DFR) technique achieves state-of-the-art group robustness via simple last-layer retraining, but it requires held-ou…

2023

UEQMS: UMAP Embedded Quick Mean Shift Algorithm for High Dimensional Clustering

AAAI 2023technical

The mean shift algorithm is a simple yet very effective clustering method widely used for image and video segmentation as well as other exploratory data analysis applications. Recently, a new algorithm called MeanShift++ (MS++) for low-dimensional clustering was proposed with a speedup of 4000 times…

Cited by 5SourcePDFScholar
2022

FedClean: A Defense Mechanism against Parameter Poisoning Attacks in Federated Learning

ICASSP 2022accepted

In Federated learning (FL) systems, a centralized entity (server), instead of access to the training data, has access to model parameter updates computed by each participant independently and based solely on their samples. Unfortunately, FL is susceptible to model poisoning attacks, in which malicio…

Cited by 0SourceScholar
2022

GridShift: A Faster Mode-Seeking Algorithm for Image Segmentation and Object Tracking

CVPR 2022oral

In machine learning, MeanShift is one of the popular clustering algorithms. It iteratively moves each data point to the weighted mean of its neighborhood data points. The computational cost required for finding neighborhood data points for each one is quadratic to the number of data points. Therefor…

Cited by 12PDFcodeScholar
2022

Master of All: Simultaneous Generalization of Urban-Scene Segmentation to All Adverse Weather Conditions

ECCV 2022poster

"Computer vision systems for autonomous navigation must generalize well in adverse weather and illumination conditions expected in the real world. However, semantic segmentation of images captured in such conditions remains a challenging task for current state-of-the-art (\sota) methods trained on b…

Cited by 14SourcePDFScholar
2021

A Scale Invariant Measure of Flatness for Deep Network Minima

ICASSP 2021accepted

It has been empirically observed that the flatness of minima obtained from training deep networks seems to correlate with better generalization. However, for deep networks with positively homogeneous activations, most measures of flatness are not invariant to rescaling of the network parameters. Thi…

Cited by 0SourceScholar
2021

Adaptive Contention Window Design Using Deep Q-Learning

ICASSP 2021accepted

We study the problem of adaptive contention window (CW) design for random-access wireless networks. More precisely, our goal is to design an intelligent node that can dynamically adapt its minimum CW (MCW) parameter to maximize a network-level utility knowing neither the MCWs of other nodes nor how…

Cited by 0SourceScholar
2021

Implicit rate-constrained optimization of non-decomposable objectives

ICML 2021spotlight

We consider a popular family of constrained optimization problems arising in machine learning that involve optimizing a non-decomposable evaluation metric with a certain thresholded form, while constraining another metric of interest. Examples of such problems include optimizing false negative rate…

2021

Retrieve in Style: Unsupervised Facial Feature Transfer and Retrieval

ICCV 2021poster

We present Retrieve in Style (RIS), an unsupervised framework for facial feature transfer and retrieval on real images. Recent work shows capabilities of transferring local facial features by capitalizing on the disentanglement property of the StyleGAN latent space. RIS improves existing art on the…

Cited by 31PDFcodeScholar
2021

Score-Based Generative Modeling through Stochastic Differential Equations

ICLR 2021oral

Creating noise from data is easy; creating data from noise is generative modeling. We present a stochastic differential equation (SDE) that smoothly transforms a complex data distribution to a known prior distribution by slowly injecting noise, and a corresponding reverse-time SDE that transforms th…

2020

Weakly Supervised Disentanglement with Guarantees

ICLR 2020poster

Learning disentangled representations that correspond to factors of variation in real-world data is critical to interpretable and human-controllable machine learning. Recently, concerns about the viability of learning disentangled representations in a purely unsupervised manner has spurred a shift t…

Cited by 169SourcecodeScholar
2019

SpotTune: Transfer Learning Through Adaptive Fine-Tuning

CVPR 2019poster

Transfer learning, which allows a source task to affect the inductive bias of the target task, is widely used in computer vision. The typical way of conducting transfer learning with deep neural networks is to fine-tune a model pretrained on the source task using data from the target task. In this p…

Cited by 640PDFScholar
2018

BlockDrop: Dynamic Inference Paths in Residual Networks

CVPR 2018poster

Very deep convolutional neural networks offer excellent recognition results, yet their computational expense limits their impact for many real-world applications. We introduce BlockDrop, an approach that learns to dynamically choose which layers of a deep network to execute during inference so as t…

2018

Co-regularized Alignment for Unsupervised Domain Adaptation

NeurIPS 2018poster

Deep neural networks, trained with large amount of labeled data, can fail to generalize well when tested with examples from a target domain whose distribution differs from the training data distribution, referred as the source domain. It can be expensive or even infeasible to obtain required amount…

2018

Delta-encoder: an effective sample synthesis method for few-shot object recognition

NeurIPS 2018spotlight

Learning to classify new categories based on just one or a few examples is a long-standing challenge in modern computer vision. In this work, we propose a simple yet effective method for few-shot (and one-shot) object recognition. Our approach is based on a modified auto-encoder, denoted delta-encod…

2018

Variational Inference of Disentangled Latent Concepts from Unlabeled Observations

ICLR 2018poster

Disentangled representations, where the higher level data generative factors are reflected in disjoint latent dimensions, offer several benefits such as ease of deriving invariant representations, transferability to other tasks, interpretability, etc. We consider the problem of unsupervised learning…

Cited by 607SourcePDFScholar
2017

Fully-Adaptive Feature Sharing in Multi-Task Networks With Applications in Person Attribute Classification

CVPR 2017spotlight

Multi-task learning aims to improve generalization performance of multiple prediction tasks by appropriately sharing relevant information across them. In the context of deep neural networks, this idea is often realized by hand-designed network architectures with layers that are shared across tasks a…

Cited by 509PDFcodeScholar
2017

Local Group Invariant Representations via Orbit Embeddings

AISTATS 2017poster

Invariance to nuisance transformations is one of the desirable properties of effective representations. We consider transformations that form a group and propose an approach based on kernel methods to derive local group invariant representations. Locality is achieved by defining a suitable probabili…

Cited by 40SourcePDFScholar
2017

S3Pool: Pooling With Stochastic Spatial Sampling

CVPR 2017poster

Feature pooling layers (e.g., max pooling) in convolutional neural networks (CNNs) serve the dual purpose of providing increasingly abstract representations as well as yielding computational savings in subsequent convolutional layers. We view the pooling operation in CNNs as a two step procedure: fi…

Cited by 106PDFcodeScholar
2017

Semi-supervised Learning with GANs: Manifold Invariance with Improved Inference

NeurIPS 2017poster

Semi-supervised learning methods using Generative adversarial networks (GANs) have shown promising empirical success recently. Most of these methods use a shared discriminator/classifier which discriminates real examples from fake while also predicting the class label. Motivated by the ability of th…

Cited by 200SourcePDFScholar
2016

Scalable Exemplar Clustering and Facility Location via Augmented Block Coordinate Descent with Column Generation

AISTATS 2016poster

In recent years exemplar clustering has become a popular tool for applications in document and video summarization, active learning, and clustering with general similarity, where cluster centroids are required to be a subset of the data samples rather than their linear combinations. The problem is…

Cited by 11SourcePDFScholar