← Search

Junkai Wu

3 accepted papers

2026

SightSound-R1: Cross-Modal Reasoning Distillation from Vision to Audio Language Models

ICASSP 2026poster

While large audio-language models (LALMs) have demonstrated state-of-the-art audio understanding, their reasoning capability in complex soundscapes still falls behind large vision-language models (LVLMs). Compared to the visual domain, one bottleneck is the lack of large-scale chain-of-thought audio…

Cited by 0SourcePDFScholar
2023

Listen, Decipher and Sign: Toward Unsupervised Speech-to-Sign Language Recognition

ACL 2023findings

Existing supervised sign language recognition systems rely on an abundance of well-annotated data. Instead, an unsupervised speech-to-sign language recognition (SSR-U) system learns to translate between spoken and sign languages by observing only non-parallel speech and sign-language corpora. We pro…