ICASSP 2025accepted0 citations

Large Language Models Are Efficient Learners as Zero-Shot Speech Translators

Chenxuan Liu, Liping Chen, Peiwang Tang, Weitai Zhang, Xiaoxi Li, Sreyan Ghosh, Zhongyi Ye, Mingjia Yu

Abstract

Significant progress has recently been made in combining Speech Foundation Models (SFMs) and Large Language Models (LLMs) into a unified model to tackle Speech-to-Text Translation (ST) tasks. However, fine-tuning LLMs to adapt to specific downstream tasks requires substantial resources, which is often infeasible. Therefore, this study proposes using Chain-of-Thought (CoT)-based LLMs to perform error correction on Automatic Speech Recognition results followed by translation into the target language. This approach combines SFMs and LLMs in a lightweight manner without fine-tuning or large parallel corpora. Additionally, through our proposed Translation Graph CoT (TGCoT), which involves iterative feedback and Back Translation, the model can self-check when errors are detected, effectively reducing error accumulation during the multi-step CoT reasoning process and thereby improving translation accuracy. Finally, extensive experiments across languages and models demonstrate the superiority and robustness of the proposed method. The results show that the proposed approach better unleashes LLM capabilities and adapts to downstream ST tasks with minimal resources.

BibTeX
@inproceedings{icassp2025_largelanguagemod,
  title = {Large Language Models Are Efficient Learners as Zero-Shot Speech Translators},
  author = {Chenxuan Liu and Liping Chen and Peiwang Tang and Weitai Zhang and Xiaoxi Li and Sreyan Ghosh and Zhongyi Ye and Mingjia Yu},
  booktitle = {ICASSP 2025},
  year = {2025}
}
Large Language Models Are Efficient Learners as Zero-Shot Speech Translators · ICASSP 2025