← Search

Xulin Fan

4 accepted papers

2026

IN-SYNC: ADAPTATION OF SPEECH AWARE LARGE LANGUAGE MODELS FOR ASR WITH WORD LEVEL TIMESTAMP PREDICTIONS

ICASSP 2026oral

Recent advances in speech-aware language models have coupled strong acoustic encoders with large language models, enabling systems that move beyond transcription to produce richer outputs. Among these, word-level timestamp prediction is critical for applications such as captioning, media search, and…

Cited by 0SourcePDFScholar
2025

ArrayDPS: Unsupervised Blind Speech Separation with a Diffusion Prior

ICML 2025poster

Blind Speech Separation (BSS) aims to separate multiple speech sources from audio mixtures recorded by a microphone array. The problem is challenging because it is a blind inverse problem, i.e., the microphone array geometry, the room impulse response (RIR), and the speech sources, are all unknown.…

2023

Dual-Path Cross-Modal Attention for Better Audio-Visual Speech Extraction

ICASSP 2023accepted

Audiovisual target speaker extraction is the task of separating, from an audio mixture, the speaker whose face is visible in an accompanying video. Published approaches typically upsample the video or downsample the audio, then fuse the two streams using concatenation, multiplication, or cross-modal…

Cited by 0SourceScholar
2023

Listen, Decipher and Sign: Toward Unsupervised Speech-to-Sign Language Recognition

ACL 2023findings

Existing supervised sign language recognition systems rely on an abundance of well-annotated data. Instead, an unsupervised speech-to-sign language recognition (SSR-U) system learns to translate between spoken and sign languages by observing only non-parallel speech and sign-language corpora. We pro…