← Search

Manoj Kumar

18 accepted papers

2026

Adapting to Evolving Graphs: A Scalable Framework for Dynamic Coarsening

ICML 2026poster

Graph coarsening is a fundamental dimensionality reduction technique for scaling large graphs while preserving structural and feature information. However, most existing coarsening methods are designed for static graphs and do not extend well to dynamic settings where nodes, edges, and connectivity …

Cited by 0SourceScholar
2025

Certifying Counterfactual Bias in LLMs

ICLR 2025poster

Large Language Models (LLMs) can produce biased responses that can cause representational harms. However, conventional studies are insufficient to thoroughly evaluate biases across LLM responses for different demographic groups (a.k.a. counterfactual bias), as they do not scale to large number of in…

Cited by 0SourcePDFScholar
2025

On Localizing and Deleting Toxic Memories in Large Language Models

NAACL 2025findings

Warning: This paper contains offensive language.Ensuring that large language models (LLMs) do not generate harmful text is critical for their safe deployment. A common failure mode involves producing toxic responses to otherwise innocuous prompts. While various detoxification methods have been propo…

Cited by 0SourcePDFScholar
2024

Frozen Feature Augmentation for Few-Shot Image Classification

CVPR 2024poster

Training a linear classifier or lightweight model on top of pretrained vision model outputs so-called 'frozen features' leads to impressive performance on a number of downstream few-shot tasks. Currently frozen features are not modified during training. On the other hand when networks are trained di…

Cited by 9SourcePDFScholar
2024

Optimization Framework for Semi-supervised Attributed Graph Coarsening

UAI 2024poster

In data-intensive applications, graphs serve as foundational structures across various domains. However, the increasing size of datasets poses significant challenges to performing downstream tasks. To address this problem, techniques such as graph coarsening, condensation, and summarization have bee…

Cited by 1SourcePDFScholar
2023

Image Captioners Are Scalable Vision Learners Too

NeurIPS 2023oral

Contrastive pretraining on image-text pairs from the web is one of the most popular large-scale pretraining strategies for vision backbones, especially in the context of large multimodal models. At the same time, image captioning on this type of data is commonly considered an inferior pretraining st…

2023

Scaling Vision Transformers to 22 Billion Parameters

ICML 2023oral

The scaling of Transformers has driven breakthrough capabilities for language models. At present, the largest large language models (LLMs) contain upwards of 100B parameters. Vision Transformers (ViT) have introduced the same architecture to image and video modelling, but these have not yet been suc…

Cited by 650SourcePDFScholar
2022

Controlled Data Generation via Insertion Operations for NLU

NAACL 2022industry

Use of synthetic data is rapidly emerging as a realistic alternative to manually annotating live traffic for industry-scale model building. Manual data annotation is slow, expensive and not preferred for meeting customer privacy expectations. Further, commercial natural language applications are req…

Cited by 6SourcePDFScholar
2022

Improving Large-Scale Conversational Assistants using Model Interpretation based Training Sample Selection

EMNLP 2022industry

This paper presents an approach to identify samples from live traffic where the customer implicitly communicated satisfaction with Alexa’s responses, by leveraging interpretations of model behavior. Such customer signals are noisy and adding a large number of samples from live traffic to training se…

Cited by 2SourcePDFScholar
2022

Unsupervised training data re-weighting for natural language understanding with local distribution approximation

EMNLP 2022industry

One of the major challenges of training Natural Language Understanding (NLU) production models lies in the discrepancy between the distributions of the offline training data and of the online live data, due to, e.g., biased sampling scheme, cyclic seasonality shifts, annotated training data coming f…

Cited by 2SourcePDFScholar
2020

Learning Domain Invariant Representations for Child-Adult Classification from Speech

ICASSP 2020accepted

Diagnostic procedures for ASD (autism spectrum disorder) involve semi-naturalistic interactions between the child and a clinician. Computational methods to analyze these sessions require an end-to-end speech and language processing pipeline that goes from raw audio to clinically-meaningful behaviora…

Cited by 0SourceScholar
2020

Meta-Learning for Robust Child-Adult Classification from Speech

ICASSP 2020accepted

Computational modeling of naturalistic conversations in clinical applications has seen growing interest in the past decade. An important use-case involves child-adult interactions within the autism diagnosis and intervention domain. In this paper, we address a specific sub-problem of speaker diariza…

Cited by 0SourceScholar
2020

Speaker Diarization Using Latent Space Clustering in Generative Adversarial Network

ICASSP 2020accepted

In this work, we propose deep latent space clustering for speaker diarization using generative adversarial network (GAN) back-projection with the help of an encoder network. The proposed diarization system is trained jointly with GAN loss, latent variable recovery loss, and a clustering-specific los…

Cited by 0SourceScholar
2020

VideoFlow: A Conditional Flow-Based Model for Stochastic Video Generation

ICLR 2020poster

Generative models that can model and predict sequences of future events can, in principle, learn to capture complex real-world phenomena, such as physical interactions. However, a central challenge in video prediction is that the future is highly uncertain: a sequence of past observations of events…

Cited by 123SourcecodeScholar
2018

Improving Semi-Supervised Classification for Low-Resource Speech Interaction Applications

ICASSP 2018accepted

We propose a semi-supervised learning method to improve classification performance in scenarios with limited labeled data. We employ adaptation strategies such as entropy-filtering and self-training, and show that our method achieves up to 17.2% relative improvement in UAR for a multi-class problem.…

Cited by 0SourceScholar