← Search

Arvind Krishna Sridhar

3 accepted papers

2025

Enhancing Temporal Understanding in Audio Question Answering for Large Audio Language Models

NAACL 2025industry

The Audio Question Answering (AQA) task includes audio event classification, audio captioning, and open-ended reasoning. Recently, AQA has garnered attention due to the advent of Large Audio Language Models (LALMs). Current literature focuses on constructing LALMs by integrating audio encoders with…

Cited by 3SourcePDFScholar
2024

Parameter Efficient Audio Captioning with Faithful Guidance Using Audio-Text Shared Latent Representation

ICASSP 2024accepted

There has been significant research on developing pretrained transformer architectures for multimodal-to-text generation tasks. Albeit performance improvements, such models frequently suffer from hallucination and large memory footprint making them challenging to deploy on edge devices. In this pape…

Cited by 0SourceScholar
2021

HypoGen: Hyperbole Generation with Commonsense and Counterfactual Knowledge

EMNLP 2021finding

A hyperbole is an intentional and creative exaggeration not to be taken literally. Despite its ubiquity in daily life, the computational explorations of hyperboles are scarce. In this paper, we tackle the under-explored and challenging task: sentence-level hyperbole generation. We start with a repre…