2024
Cross-Modal Alignment for End-to-End Spoken Language Understanding Based on Momentum Contrastive Learning
ICASSP 2024accepted
The end-to-end spoken language understanding system extracts the semantic intent directly from an input speech. It effectively avoids problems such as semantic drift in traditional cascade models. However, the lack of semantically labeled speech data makes the model training process diffi-cult. Seve…