← Search

Rita Singh

37 accepted papers

2025

ADIFF: Explaining audio difference using natural language

ICLR 2025spotlight

Understanding and explaining differences between audio recordings is crucial for fields like audio forensics, quality assessment, and audio generation. This involves identifying and describing audio events, acoustic scenes, signal characteristics, and their emotional impact on listeners. This paper…

2025

Audio Entailment: Assessing Deductive Reasoning for Audio Understanding

AAAI 2025technical

Recent literature uses language to build foundation models for audio. These Audio-Language Models (ALMs) are trained on a vast number of audio-text pairs and show remarkable performance in tasks including Text-to-Audio Retrieval, Captioning, and Question Answering. However, their ability to engage i…

2025

CAARMA: Class Augmentation with Adversarial Mixup Regularization

EMNLP 2025

Speaker verification is a typical zero-shot learning task, where inference of unseen classes is performed by comparing embeddings of test instances to known examples. The models performing inference must hence naturally generate embeddings that cluster same-class instances compactly, while maintaini

2025

Lost in Transcription, Found in Distribution Shift: Demystifying Hallucination in Speech Foundation Models

ACL 2025finding

Speech foundation models trained at a massive scale, both in terms of model and data size, result in robust systems capable of performing multiple speech tasks, including automatic speech recognition (ASR). These models transcend language and domain barriers, yet effectively measuring their performa…

2025

PhoniTale: Phonologically Grounded Mnemonic Generation for Typologically Distant Language Pairs

EMNLP 2025

Vocabulary acquisition poses a significant challenge for second-language (L2) learners, especially when learning typologically distant languages such as English and Korean, where phonological and structural mismatches complicate vocabulary learning. Recently, large language models (LLMs) have been u

2025

SVeritas: Benchmark for Robust Speaker Verification under Diverse Conditions

EMNLP 2025

Speaker verification (SV) models are increasingly integrated into security, personalization, and access control systems, yet their robustness to many real-world challenges remains inadequately benchmarked. Real-world systems can face diverse conditions, some naturally occurring, and others that may

2024

A General Framework for Learning from Weak Supervision

ICML 2024poster

Weakly supervised learning generally faces challenges in applicability to various scenarios with diverse weak supervision and in scalability due to the complexity of existing algorithms, thereby hindering the practical deployment. This paper introduces a general framework for learning from weak supe…

2024

Completing Visual Objects via Bridging Generation and Segmentation

ICML 2024poster

This paper presents a novel approach to object completion, with the primary goal of reconstructing a complete object from its partially visible components. Our method, named MaskComp, delineates the completion process through iterative stages of generation and segmentation. In each iteration, the ob…

Cited by 6SourcePDFScholar
2024

Imprecise Label Learning: A Unified Framework for Learning with Various Imprecise Label Configurations

NeurIPS 2024poster

Learning with reduced labeling standards, such as noisy label, partial label, and supplementary unlabeled data, which we generically refer to as imprecise label, is a commonplace challenge in machine learning tasks. Previous methods tend to propose specific designs for every emerging imprecise label…

2024

Prompting Audios Using Acoustic Properties for Emotion Representation

ICASSP 2024accepted

Emotions lie on a continuum, but current models treat emotions as a finite valued discrete variable. This representation does not capture the diversity in the expression of emotion. To better represent emotions we propose the use of natural language descriptions (or prompts). In this work, we addres…

Cited by 0SourceScholar
2024

QDFormer: Towards Robust Audiovisual Segmentation in Complex Environments with Quantization-based Semantic Decomposition

CVPR 2024poster

Audiovisual segmentation (AVS) is a challenging task that aims to segment visual objects in videos according to their associated acoustic cues. With multiple sound sources and background disturbances involved establishing robust correspondences between audio and visual contents poses unique challeng…

2024

R-BASS : Relevance-aided Block-wise Adaptation for Speech Summarization

NAACL 2024findings

End-to-end speech summarization on long recordings is challenging because of the high computational cost. Block-wise Adaptation for Speech Summarization (BASS) summarizes arbitrarily long sequences by sequentially processing abutting chunks of audio. Despite the benefits of BASS, it has higher compu…

Cited by 0SourcePDFScholar
2024

R^2-Bench: Benchmarking the Robustness of Referring Perception Models under Perturbations

ECCV 2024poster

"Referring perception, which aims at grounding visual objects with multimodal referring guidance, is essential for bridging the gap between humans, who provide instructions, and the environment where intelligent systems perceive. Despite progress in this field, the robustness of referring perception…

Cited by 3SourcePDFScholar
2024

Training Audio Captioning Models without Audio

ICASSP 2024accepted

Automated Audio Captioning (AAC) is the task of generating natural language descriptions given an audio stream. A typical AAC system requires manually curated training data of audio segments and corresponding text caption annotations. The creation of these audio-caption pairs is costly, resulting in…

Cited by 0SourceScholar
2023

PaintSeg: Painting Pixels for Training-free Segmentation

NeurIPS 2023poster

The paper introduces PaintSeg, a new unsupervised method for segmenting objects without any training. We propose an adversarial masked contrastive painting (AMCP) process, which creates a contrast between the original image and a painted image in which a masked area is painted using off-the-shelf ge…

2023

Pairwise Similarity Learning is SimPLE

ICCV 2023poster

In this paper, we focus on a general yet important learning problem, pairwise similarity learning (PSL). PSL subsumes a wide range of important applications, such as open-set face recognition, speaker verification, image retrieval and person re-identification. The goal of PSL is to learn a pairwise…

Cited by 10PDFcodeScholar
2023

Pengi: An Audio Language Model for Audio Tasks

NeurIPS 2023poster

In the domain of audio processing, Transfer Learning has facilitated the rise of Self-Supervised Learning and Zero-Shot Learning techniques. These approaches have led to the development of versatile models capable of tackling a wide array of tasks, while delivering state-of-the-art performance. Howe…

2023

Token Prediction as Implicit Classification to Identify LLM-Generated Text

EMNLP 2023short main

This paper introduces a novel approach for identifying the possible large language models (LLMs) involved in text generation. Instead of adding an additional classification layer to a base LM, we reframe the classification task as a next-token prediction task and directly fine-tune the base LM to pe…

Cited by 0SourcecodeScholar
2023

Towards Noise-Tolerant Speech-Referring Video Object Segmentation: Bridging Speech and Text

EMNLP 2023long main

Linguistic communication is prevalent in Human-Computer Interaction (HCI). Speech (spoken language) serves as a convenient yet potentially ambiguous form due to noise and accents, exposing a gap compared to text. In this study, we investigate the prominent HCI task, Referring Video Object Segmentati…

Cited by 0SourceScholar
2022

SphereFace2: Binary Classification is All You Need for Deep Face Recognition

ICLR 2022spotlight

State-of-the-art deep face recognition methods are mostly trained with a softmax-based multi-class classification framework. Despite being popular and effective, these methods still have a few shortcomings that limit empirical performance. In this paper, we start by identifying the discrepancy betwe…

Cited by 63SourcePDFScholar
2020

Speech-Based Parameter Estimation of an Asymmetric Vocal Fold Oscillation Model and its Application in Discriminating Vocal Fold Pathologies

ICASSP 2020accepted

So far, several physical models have been proposed for the study of vocal fold oscillations during phonation. The parameters of these models, such as vocal fold elasticity, resistance, etc. are traditionally determined through the observation and measurement of the vocal fold vibrations in the laryn…

Cited by 0SourceScholar
2019

Disjoint Mapping Network for Cross-modal Matching of Voices and Faces

ICLR 2019poster

We propose a novel framework, called Disjoint Mapping Network (DIMNet), for cross-modal biometric matching, in particular of voices and faces. Different from the existing methods, DIMNet does not explicitly learn the joint relationship between the modalities. Instead, DIMNet learns a shared represen…

Cited by 92SourcePDFScholar
2019

Face Reconstruction from Voice using Generative Adversarial Networks

NeurIPS 2019poster

Voice profiling aims at inferring various human parameters from their speech, e.g. gender, age, etc. In this paper, we address the challenge posed by a subtask of voice profiling - reconstructing someone's face from their voice. The task is designed to answer the question: given an audio clip spoken…

2018

A Corrective Learning Approach for Text-Independent Speaker Verification

ICASSP 2018accepted

We present a conceptually plausible approach for text-independent speaker verification (TISV) which treats speech recordings as a collection of segments providing incremental evidence. This approach, called corrective learning, gradually improves an initial prediction of speaker identity based on in…

Cited by 0SourceScholar
2016

The relationship of voice onset time and Voice Offset Time to physical age

ICASSP 2016accepted

In a speech signal, Voice Onset Time (VOT) is the period between the release of a plosive and the onset of vocal cord vibrations in the production of the following sound. Voice Offset Time (VOFT), on the other hand, is the period between the end of a voiced sound and the release of the following plo…

Cited by 0SourceScholar