← Search

Ashish Seth

15 accepted papers

2026

EgoAVU: Egocentric Audio-Visual Understanding

CVPR 2026

Understanding egocentric videos plays a vital role for embodied intelligence. Recent multi-modal large language models (MLLMs) can accept both visual and audio inputs. However, due to the challenge of obtaining text labels with coherent joint-modality information, whether MLLMs can jointly understan

Cited by 0SourcecodeScholar
2026

Exploring Audio Hallucination in Egocentric Video Understanding

ICASSP 2026oral

Egocentric videos provide a distinctive setting in which sound serves as crucial cues to understand user activities and surroundings, particularly when visual information is unstable or occluded due to continuous camera movement. State-of-the-art large audio-visual language models (AV-LLMs) can gene…

Cited by 0SourcePDFScholar
2025

Do Audio-Language Models Understand Linguistic Variations?

NAACL 2025short

Open-vocabulary audio language models (ALMs), like Contrastive Language Audio Pretraining (CLAP), represent a promising new paradigm for audio-text retrieval using natural language queries. In this paper, for the first time, we perform controlled experiments on various benchmarks to show that existi…

2025

EGOILLUSION: Benchmarking Hallucinations in Egocentric Video Understanding

EMNLP 2025

Multimodal Large Language Models (MLLMs) have demonstrated remarkable performance in complex multimodal tasks. While MLLMs excel at visual perception and reasoning in third-person and egocentric videos, they are prone to hallucinations, generating coherent yet inaccurate responses. We present EGOILL

2025

MMAU: A Massive Multi-Task Audio Understanding and Reasoning Benchmark

ICLR 2025spotlight

The ability to comprehend audio—which includes speech, non-speech sounds, and music—is crucial for AI agents to interact effectively with the world. We present MMAU, a novel benchmark designed to evaluate multimodal audio understanding models on tasks requiring expert-level knowledge and complex rea…

Cited by 25SourcePDFScholar
2025

MULTIVOX: A Benchmark for Evaluating Voice Assistants for Multimodal Interactions

EMNLP 2025

The rapid progress of Large Language Models (LLMs) has empowered omni models to act as voice assistants capable of understanding spoken dialogues. These models can process multimodal inputs beyond text, such as speech and visual data, enabling more context-aware interactions. However, current benchm

2025

PAT: Parameter-Free Audio-Text Aligner to Boost Zero-Shot Audio Classification

NAACL 2025long

Audio-Language Models (ALMs) have demonstrated remarkable performance in zero-shot audio classification. In this paper, we introduce PAT (Parameter-free Audio-Text aligner), a simple and training-free method aimed at boosting zero-shot audio classification performance of CLAP-like ALMs. To achieve t…

2024

CompA: Addressing the Gap in Compositional Reasoning in Audio-Language Models

ICLR 2024poster

A fundamental characteristic of audio is its compositional nature. Audio-language models (ALMs) trained using a contrastive approach (e.g., CLAP) that learns a shared representation between audio and language modalities have improved performance in many downstream applications, including zero-shot a…

Cited by 12SourcePDFScholar
2024

EH-MAM: Easy-to-Hard Masked Acoustic Modeling for Self-Supervised Speech Representation Learning

EMNLP 2024main

In this paper, we present EH-MAM (Easy-to-Hard adaptive Masked Acoustic Modeling), a novel self-supervised learning approach for speech representation learning. In contrast to the prior methods that use random masking schemes for Masked Acoustic Modeling (MAM), we introduce a novel selective and ada…

2024

FusDom: Combining in-Domain and Out-of-Domain Knowledge for Continuous Self-Supervised Learning

ICASSP 2024accepted

Continued pre-training (CP) offers multiple advantages, like target domain adaptation and the potential to exploit the continuous stream of unlabeled data available online. However, continued pre-training on out-of-domain distributions often leads to catastrophic forgetting of previously acquired kn…

Cited by 0SourceScholar
2024

GAMA: A Large Audio-Language Model with Advanced Audio Understanding and Complex Reasoning Abilities

EMNLP 2024main

Perceiving and understanding non-speech sounds and non-verbal speech is essential to making decisions that help us interact with our surroundings. In this paper, we propose GAMA, a novel General-purpose Large Audio-Language Model (LALM) with Advanced Audio Understanding and Complex Reasoning Abiliti…

2024

Stable Distillation: Regularizing Continued Pre-Training for Low-Resource Automatic Speech Recognition

ICASSP 2024accepted

Continued self-supervised (SSL) pre-training for adapting existing SSL models to the target domain has shown to be extremely effective for low-resource Automatic Speech Recognition (ASR). This paper proposes Stable Distillation, a simple and novel approach for SSL-based continued pre-training that b…

Cited by 0SourceScholar
2023

SLICER: Learning Universal Audio Representations Using Low-Resource Self-Supervised Pre-Training

ICASSP 2023accepted

We present a new Self-Supervised Learning (SSL) approach to pre-train encoders on unlabeled audio data that reduces the need for large amounts of labeled data for audio and speech classification. Our primary aim is to learn au-dio representations that can generalize across a large vari-ety of speech…

Cited by 0SourceScholar