← Search

Utkarsh Tyagi

16 accepted papers

2026

SciPredict: Can LLMs Predict the Outcomes of Scientific Experiments in Natural Sciences?

ICML 2026poster

Accelerating scientific discovery requires the identification of which experiments would yield the best outcomes before committing resources to costly physical validation. While existing benchmarks evaluate LLMs on scientific knowledge and reasoning, their ability to predict experimental outcomes---…

Cited by 0SourceScholar
2025

EGOILLUSION: Benchmarking Hallucinations in Egocentric Video Understanding

EMNLP 2025

Multimodal Large Language Models (MLLMs) have demonstrated remarkable performance in complex multimodal tasks. While MLLMs excel at visual perception and reasoning in third-person and egocentric videos, they are prone to hallucinations, generating coherent yet inaccurate responses. We present EGOILL

2025

MMAU: A Massive Multi-Task Audio Understanding and Reasoning Benchmark

ICLR 2025spotlight

The ability to comprehend audio—which includes speech, non-speech sounds, and music—is crucial for AI agents to interact effectively with the world. We present MMAU, a novel benchmark designed to evaluate multimodal audio understanding models on tasks requiring expert-level knowledge and complex rea…

Cited by 25SourcePDFScholar
2025

MULTIVOX: A Benchmark for Evaluating Voice Assistants for Multimodal Interactions

EMNLP 2025

The rapid progress of Large Language Models (LLMs) has empowered omni models to act as voice assistants capable of understanding spoken dialogues. These models can process multimodal inputs beyond text, such as speech and visual data, enabling more context-aware interactions. However, current benchm

2025

ProSE: Diffusion Priors for Speech Enhancement

NAACL 2025long

Speech enhancement (SE) is the fundamental task of enhancing the clarity and quality of speech in the presence of non-stationary additive noise. While deterministic deep learning models have been commonly employed for SE, recent research indicates that generative models, such as denoising diffusion…

2025

Visual Description Grounding Reduces Hallucinations and Boosts Reasoning in LVLMs

ICLR 2025poster

Large Vision-Language Models (LVLMs) often produce responses that misalign with factual information, a phenomenon known as hallucinations. While hallucinations are well-studied, the exact causes behind them remain underexplored. In this paper, we first investigate the root causes of hallucinations i…

2024

ABEX: Data Augmentation for Low-Resource NLU via Expanding Abstract Descriptions

ACL 2024long

We present ABEX, a novel and effective generative data augmentation methodology for low-resource Natural Language Understanding (NLU) tasks. ABEX is based on ABstract-and-EXpand, a novel paradigm for generating diverse forms of an input document – we first convert a document into its concise, abstra…

2024

ASPIRE: Language-Guided Data Augmentation for Improving Robustness Against Spurious Correlations

ACL 2024findings

Neural image classifiers can often learn to make predictions by overly relying on non-predictive features that are spuriously correlated with the class labels in the training data. This leads to poor performance in real-world atypical scenarios where such features are absent. This paper presents ASP…

2024

CoDa: Constrained Generation based Data Augmentation for Low-Resource NLP

NAACL 2024findings

We present CoDa (**Co**nstrained Generation based **Da**ta Augmentation), a controllable, effective, and *training-free* data augmentation technique for low-resource (data-scarce) NLP. Our approach is based on prompting off-the-shelf instruction-following Large Language Models (LLMs) for generating…

2024

CompA: Addressing the Gap in Compositional Reasoning in Audio-Language Models

ICLR 2024poster

A fundamental characteristic of audio is its compositional nature. Audio-language models (ALMs) trained using a contrastive approach (e.g., CLAP) that learns a shared representation between audio and language modalities have improved performance in many downstream applications, including zero-shot a…

Cited by 12SourcePDFScholar
2024

Do Vision-Language Models Understand Compound Nouns?

NAACL 2024short

Open-vocabulary vision-language models (VLMs) like CLIP, trained using contrastive loss, have emerged as a promising new paradigm for text-to-image retrieval. However, do VLMs understand compound nouns (CNs) (e.g., *lab coat*) as well as they understand nouns (e.g., *lab*)? We curate Compun, a novel…

2024

GAMA: A Large Audio-Language Model with Advanced Audio Understanding and Complex Reasoning Abilities

EMNLP 2024main

Perceiving and understanding non-speech sounds and non-verbal speech is essential to making decisions that help us interact with our surroundings. In this paper, we propose GAMA, a novel General-purpose Large Audio-Language Model (LALM) with Advanced Audio Understanding and Complex Reasoning Abiliti…

2023

ACLM: A Selective-Denoising based Generative Data Augmentation Approach for Low-Resource Complex NER

ACL 2023long

Complex Named Entity Recognition (NER) is the task of detecting linguistically complex named entities in low-context text. In this paper, we present ACLM Attention-map aware keyword selection for Conditional Language Model fine-tuning), a novel data augmentation approach based on conditional generat…

2023

AdVerb: Visually Guided Audio Dereverberation

ICCV 2023poster

We present AdVerb, a novel audio-visual dereverberation framework that uses visual cues in addition to the reverberant sound to estimate clean audio. Although audio-only dereverberation is a well-studied problem, our approach incorporates the complementary visual modality to perform audio dereverber…

Cited by 10PDFScholar
2023

CoSyn: Detecting Implicit Hate Speech in Online Conversations Using a Context Synergized Hyperbolic Network

EMNLP 2023long main

The tremendous growth of social media users interacting in online conversations has led to significant growth in hate speech affecting people from various demographics. Most of the prior works focus on detecting explicit hate speech, which is overt and leverages hateful phrases, with very little wor…

Cited by 0SourcecodeScholar
2023

DALE: Generative Data Augmentation for Low-Resource Legal NLP

EMNLP 2023long main

We present DALE, a novel and effective generative Data Augmentation framework for low-resource LEgal NLP. DALE addresses the challenges existing frameworks pose in generating effective data augmentations of legal documents - legal language, with its specialized vocabulary and complex semantics, morp…

Cited by 0SourcecodeScholar