ICML 2025poster32 citations

Freeze-Omni: A Smart and Low Latency Speech-to-speech Dialogue Model with Frozen LLM

Xiong Wang, Yangze Li, Chaoyou Fu, Yike Zhang, Yunhang Shen, Lei Xie, Ke Li, Xing Sun

Abstract

The GPT-4o's excellent duplex speech interaction ability has given users an impressive experience. Researchers have recently proposed several multimodal LLMs to achieve user-agent speech-to-speech conversations. In this paper, we propose a novel speech-text multimodal LLM architecture called Freeze-Omni, and our main contribution is that the speech input and output modalities can be easily connected to a textual LLM while keeping the LLM's parameters frozen throughout the training process. We effectively ensure that the intelligence of the Freeze-Omni in the speech modality is at the same level as that in the text modality of its backbone LLM while achieving low latency in the end-to-end spoken response. In addition, we also designed a method to achieve duplex dialogue ability through multitask training, giving Freeze-Omni a more natural style of dialogue ability between users and agents. In summary, Freeze-Omni holds great potential to conduct speech-to-speech dialogue based on a multimodal LLM under the condition of a frozen LLM, avoiding the catastrophic forgetting problem caused by limited data and training resources.

Speech to SpeechDuplex Dialogue ModelMultimodal Large Language Models
BibTeX
@inproceedings{
wang2025freezeomni,
title={Freeze-Omni: A Smart and Low Latency Speech-to-speech Dialogue Model with Frozen {LLM}},
author={Xiong Wang and Yangze Li and Chaoyou Fu and Yike Zhang and Yunhang Shen and Lei Xie and Ke Li and Xing Sun and Long MA},
booktitle={Forty-second International Conference on Machine Learning},
year={2025},
url={https://openreview.net/forum?id=s1EImzs5Id}
}
Freeze-Omni: A Smart and Low Latency Speech-to-speech Dialogue Model with Frozen LLM · ICML 2025