← Search

Jiliang Hu

4 accepted papers

2026

End-to-End Contrastive Language-Speech Pretraining Model for Long-Form Spoken Question Answering

AAAI 2026technical

Significant progress has been made in spoken question answering (SQA) in recent years. However, many existing methods, including large audio language models, struggle with processing long audio. Follow the success of retrieval augmented generation, a speech-related retriever shows promising in help

Cited by 0SourcePDFScholar
2025

Joint Automatic Speech Recognition And Structure Learning For Better Speech Understanding

ICASSP 2025accepted

Spoken language understanding (SLU) is a structure prediction task in the field of speech. Recently, many works on SLU that treat it as a sequence-to-sequence task have achieved great success. However, This method is not suitable for simultaneous speech recognition and understanding. In this paper,…

Cited by 0SourceScholar
2025

SongSong: A Time Phonograph for Chinese SongCi Music from Thousand of Years Away

AAAI 2025technical

Recently, there have been significant advancements in music generation. However, existing models primarily focus on creating modern pop songs, making it challenging to produce ancient music with distinct rhythms and styles, such as ancient Chinese SongCi. In this paper, we introduce SongSong, the fi…

Cited by 0SourcePDFScholar
2024

VHASR: A Multimodal Speech Recognition System With Vision Hotwords

EMNLP 2024main

The image-based multimodal automatic speech recognition (ASR) model enhances speech recognition performance by incorporating audio-related image. However, some works suggest that introducing image information to model does not help improving ASR performance. In this paper, we propose a novel approac…