EXPLORING FINE-TUNING OF LARGE AUDIO LANGUAGE MODELS FOR SPOKEN LANGUAGE UNDERSTANDING UNDER LIMITED SPEECH DATA
Large Audio Language Models (LALMs) have emerged as powerful tools for speech-related tasks but remain underexplored for fine-tuning, especially with limited speech data. To bridge this gap, we systematically examine how different fine-tuning schemes including text-only, direct mixing, and curriculu…