← Search

Xuanru Zhou

4 accepted papers

2026

K-FUNCTION: JOINT PRONUNCIATION TRANSCRIPTION AND FEEDBACK FOR EVALUATING KIDS LANGUAGE FUNCTION

ICASSP 2026poster

Evaluating young children's language is challenging for automatic speech recognizers due to high-pitched voices, prolonged sounds, and limited data. We introduce K-Function, a framework that combines accurate sub-word transcription with objective, Large Language Model (LLM)-driven scoring. Its core,…

Cited by 0SourcePDFScholar
2026

Speech World Model: Causal State–Action Planning with Explicit Reasoning for Speech

ICLR 2026poster

Current speech-language models (SLMs) typically use a cascade of speech encoder and large language model, treating speech understanding as a single black box. They analyze the content of speech well but reason weakly about other aspects, especially under sparse supervision. Thus, we argue for explic…

Cited by 0SourceScholar
2026

Unlocking Strong Supervision: A Data-Centric Study of General-Purpose Audio Pre-Training Methods

CVPR 2026

Current audio pre-training seeks to learn unified representations for broad audio understanding tasks, but it remains fragmented and is fundamentally bottlenecked by its reliance on weak, noisy, and scale-limited labels. Drawing lessons from vision's foundational pre-training blueprint, we argue tha

Cited by 0SourceScholar
2024

SSDM: Scalable Speech Dysfluency Modeling

NeurIPS 2024poster

Speech dysfluency modeling is the core module for spoken language learning, and speech therapy. However, there are three challenges. First, current state-of-the-art solutions~~\cite{lian2023unconstrained-udm, lian-anumanchipalli-2024-towards-hudm} suffer from poor scalability. Second, there is a lac…