← Search

Ramaneswaran Selvakumar

7 accepted papers

2025

Do Audio-Language Models Understand Linguistic Variations?

NAACL 2025short

Open-vocabulary audio language models (ALMs), like Contrastive Language Audio Pretraining (CLAP), represent a promising new paradigm for audio-text retrieval using natural language queries. In this paper, for the first time, we perform controlled experiments on various benchmarks to show that existi…

2025

EGOILLUSION: Benchmarking Hallucinations in Egocentric Video Understanding

EMNLP 2025

Multimodal Large Language Models (MLLMs) have demonstrated remarkable performance in complex multimodal tasks. While MLLMs excel at visual perception and reasoning in third-person and egocentric videos, they are prone to hallucinations, generating coherent yet inaccurate responses. We present EGOILL

2025

MMAU: A Massive Multi-Task Audio Understanding and Reasoning Benchmark

ICLR 2025spotlight

The ability to comprehend audio—which includes speech, non-speech sounds, and music—is crucial for AI agents to interact effectively with the world. We present MMAU, a novel benchmark designed to evaluate multimodal audio understanding models on tasks requiring expert-level knowledge and complex rea…

Cited by 25SourcePDFScholar
2025

MULTIVOX: A Benchmark for Evaluating Voice Assistants for Multimodal Interactions

EMNLP 2025

The rapid progress of Large Language Models (LLMs) has empowered omni models to act as voice assistants capable of understanding spoken dialogues. These models can process multimodal inputs beyond text, such as speech and visual data, enabling more context-aware interactions. However, current benchm

2025

PAT: Parameter-Free Audio-Text Aligner to Boost Zero-Shot Audio Classification

NAACL 2025long

Audio-Language Models (ALMs) have demonstrated remarkable performance in zero-shot audio classification. In this paper, we introduce PAT (Parameter-free Audio-Text aligner), a simple and training-free method aimed at boosting zero-shot audio classification performance of CLAP-like ALMs. To achieve t…

2024

EH-MAM: Easy-to-Hard Masked Acoustic Modeling for Self-Supervised Speech Representation Learning

EMNLP 2024main

In this paper, we present EH-MAM (Easy-to-Hard adaptive Masked Acoustic Modeling), a novel self-supervised learning approach for speech representation learning. In contrast to the prior methods that use random masking schemes for Masked Acoustic Modeling (MAM), we introduce a novel selective and ada…