ICASSP 2025accepted0 citations

Can RAG-Driven Enhancements Amplify Audio LLMs for Low-Resource Languages?

Bikash Dutta, Rishabh Ranjan, Akshat Jain, Richa Singh, Mayank Vatsa

Abstract

The proliferation of Large Language Models (LLMs) has transformed Natural Language Processing (NLP), yet their development has largely overlooked low-resource languages. This paper addresses this disparity by evaluating three prominent Large Audio Language Models (LALMs) – LTU-AS, GAMA, and Pengi – across tasks like Automatic Speech Recognition (ASR), Audio Question Answering (AQA), and audio classification tasks in Hindi and code-mixed Hindi-English (aka Hinglish). We also explore the potential of Retrieval-Augmented Generation (RAG) to boost LALM performance in these low-resource settings. Our findings highlight significant performance discrepancies, with LALMs performing well in audio classification but struggling with ASR and AQA. While RAG shows potential, especially for audio classification, its impact is inconsistent across tasks. This work offers critical insights into the challenges of using LALMs for low-resource languages and provides a foundation for developing more inclusive and adaptable AI systems for complex multilingual tasks.

BibTeX
@inproceedings{icassp2025_canragdrivenenha,
  title = {Can RAG-Driven Enhancements Amplify Audio LLMs for Low-Resource Languages?},
  author = {Bikash Dutta and Rishabh Ranjan and Akshat Jain and Richa Singh and Mayank Vatsa},
  booktitle = {ICASSP 2025},
  year = {2025}
}