Turn-Taking and Backchannel Prediction with Acoustic and Large Language Model Fusion
Jinhan Wang, Long Chen, Aparna Khare, Anirudh Raju, Pranav Dheram, Di He, Minhua Wu, Andreas Stolcke
Abstract
We propose an approach for continuous prediction of turn-taking and backchanneling locations in spoken dialogue by fusing a neural acoustic model with a large language model (LLM). Experiments on the Switchboard human-human conversation dataset demonstrate that our approach consistently outperforms the baseline models with single modality. We also develop a novel multi-task instruction fine-tuning strategy to further benefit from LLM-encoded knowledge for understanding the tasks and conversational contexts, leading to additional improvements. Our approach demonstrates the potential of combined LLMs and acoustic models for a more natural and conversational interaction between humans and speech-enabled AI agents.
BibTeX
@inproceedings{icassp2024_turntakingandbac,
title = {Turn-Taking and Backchannel Prediction with Acoustic and Large Language Model Fusion},
author = {Jinhan Wang and Long Chen and Aparna Khare and Anirudh Raju and Pranav Dheram and Di He and Minhua Wu and Andreas Stolcke and Venkatesh Ravichandran},
booktitle = {ICASSP 2024},
year = {2024}
}