← Search

Jiachen Lian

7 accepted papers

2026

FULL-DUPLEX-BENCH V1.5: EVALUATING OVERLAP HANDLING FOR FULL-DUPLEX SPEECH MODELS

ICASSP 2026poster

Full-duplex spoken dialogue systems promise to transform human-machine interaction from a rigid, turn-based protocol into a fluid, natural conversation. However, the central challenge to realizing this vision, managing overlapping speech, remains critically under-evaluated. We introduce Full-Duplex-…

Cited by 0SourcePDFScholar
2026

K-FUNCTION: JOINT PRONUNCIATION TRANSCRIPTION AND FEEDBACK FOR EVALUATING KIDS LANGUAGE FUNCTION

ICASSP 2026poster

Evaluating young children's language is challenging for automatic speech recognizers due to high-pitched voices, prolonged sounds, and limited data. We introduce K-Function, a framework that combines accurate sub-word transcription with objective, Large Language Model (LLM)-driven scoring. Its core,…

Cited by 0SourcePDFScholar
2026

Speech World Model: Causal State–Action Planning with Explicit Reasoning for Speech

ICLR 2026poster

Current speech-language models (SLMs) typically use a cascade of speech encoder and large language model, treating speech understanding as a single black box. They analyze the content of speech well but reason weakly about other aspects, especially under sparse supervision. Thus, we argue for explic…

Cited by 0SourceScholar
2025

Seeing is Believing: Emotion-Aware Audio-Visual Language Modeling for Expressive Speech Generation

EMNLP 2025

We present an Audio-Visual Language Model (AVLM) for expressive speech generation by integrating full-face visual cues into a pre-trained expressive speech model. We explore multiple visual encoders and multimodal fusion strategies during pre-training to identify the most effective integration appro

2024

SSDM: Scalable Speech Dysfluency Modeling

NeurIPS 2024poster

Speech dysfluency modeling is the core module for spoken language learning, and speech therapy. However, there are three challenges. First, current state-of-the-art solutions~~\cite{lian2023unconstrained-udm, lian-anumanchipalli-2024-towards-hudm} suffer from poor scalability. Second, there is a lac…

2023

Articulatory Representation Learning via Joint Factor Analysis and Neural Matrix Factorization

ICASSP 2023accepted

Articulatory representation learning is the fundamental research in modeling neural speech production system. Our previous work has established a deep paradigm to decompose the articulatory kinematics data into gestures, which explicitly model the phonological and linguistic structure encoded with h…

Cited by 0SourceScholar
2022

Robust Disentangled Variational Speech Representation Learning for Zero-Shot Voice Conversion

ICASSP 2022accepted

Traditional studies on voice conversion (VC) have made progress with parallel training data and known speakers. Good voice conversion quality is obtained by exploring better alignment modules or expressive mapping functions. In this study, we investigate zero-shot VC from a novel perspective of self…

Cited by 0SourceScholar