End-to-End Contrastive Language-Speech Pretraining Model for Long-Form Spoken Question Answering
Significant progress has been made in spoken question answering (SQA) in recent years. However, many existing methods, including large audio language models, struggle with processing long audio. Follow the success of retrieval augmented generation, a speech-related retriever shows promising in help