← Search

Erik Visser

4 accepted papers

2025

Enhancing Temporal Understanding in Audio Question Answering for Large Audio Language Models

NAACL 2025industry

The Audio Question Answering (AQA) task includes audio event classification, audio captioning, and open-ended reasoning. Recently, AQA has garnered attention due to the advent of Large Audio Language Models (LALMs). Current literature focuses on constructing LALMs by integrating audio encoders with…

Cited by 3SourcePDFScholar
2024

Parameter Efficient Audio Captioning with Faithful Guidance Using Audio-Text Shared Latent Representation

ICASSP 2024accepted

There has been significant research on developing pretrained transformer architectures for multimodal-to-text generation tasks. Albeit performance improvements, such models frequently suffer from hallucination and large memory footprint making them challenging to deploy on edge devices. In this pape…

Cited by 0SourceScholar
2022

Multi-Task Voice Activated Framework Using Self-Supervised Learning

ICASSP 2022accepted

Self-supervised learning methods such as wav2vec 2.0 have shown promising results in learning speech representations from unlabelled and untranscribed speech data that are useful for speech recognition. Since these representations are learned without any task-specific supervision, they can also be u…

Cited by 0SourceScholar