Bridging Modality Gap with Large Speech and Language Models for End-to-End Speech-to-Text Translation
End-to-end speech-to-text translation (E2E ST) has increasingly aroused interest and attention recently, attempting to address the problem of data scarcity and modeling burden. Several attempts exploring the combination of Large Speech and Language Models into a unified model to improve E2E ST are c…