Adversarial Speech-Text Pre-Training for Speech Translation
Chenxuan Liu, Liping Chen, Weitai Zhang, Xiaoxi Li, Peiwang Tang, Mingjia Yu, Sreyan Ghosh, Zhongyi Ye
Abstract
Large-scale pre-training has been shown to benefit speech translation tasks. However, existing multimodal pre-training efforts rely on parallel corpora for semantic alignment, potentially limiting performance to the scale of available data and causing data imbalance. Hence, we propose an adversarial speech-text pre-training (AST) scheme for speech translation. This scheme aligns the feature distributions of speech and text modalities instead of enforcing semantic alignment based on parallel corpora. Specifically, we introduced a dual-stream mechanism that bridges speech and text modalities through speech, text, and shared encoders. In addition, we designed an adversarial bridging method that focuses on the differences in feature distributions between speech and text. Leveraging a discriminator and hidden state-level swapping strategy emphasizes semantic information in speech representations, while avoiding the limitations imposed by the scale of parallel corpora. We applied AST to both end-to-end speech translation and large model architectures. Experimental results on the IWSLT test sets demonstrate that AST improved the performance of speech translation models and is compatible with large language models.
BibTeX
@inproceedings{icassp2025_adversarialspeec,
title = {Adversarial Speech-Text Pre-Training for Speech Translation},
author = {Chenxuan Liu and Liping Chen and Weitai Zhang and Xiaoxi Li and Peiwang Tang and Mingjia Yu and Sreyan Ghosh and Zhongyi Ye},
booktitle = {ICASSP 2025},
year = {2025}
}