← Search

Qiaolin Wang

2 accepted papers

2026

SightSound-R1: Cross-Modal Reasoning Distillation from Vision to Audio Language Models

ICASSP 2026poster

While large audio-language models (LALMs) have demonstrated state-of-the-art audio understanding, their reasoning capability in complex soundscapes still falls behind large vision-language models (LVLMs). Compared to the visual domain, one bottleneck is the lack of large-scale chain-of-thought audio…

Cited by 0SourcePDFScholar
2025

Layer-wise Minimal Pair Probing Reveals Contextual Grammatical-Conceptual Hierarchy in Speech Representations

EMNLP 2025

Transformer-based speech language models (SLMs) have significantly improved neural speech recognition and understanding. While existing research has examined how well SLMs encode shallow acoustic and phonetic features, the extent to which SLMs encode nuanced syntactic and conceptual features remains

Cited by 0SourcePDFScholar