← Search

Youjun Chen

5 accepted papers

2026

EmotionThinker: Prosody-Aware Reinforcement Learning for Explainable Speech Emotion Reasoning

ICLR 2026oral

Emotional information in speech plays a unique role in multimodal perception. However, current Speech Large Language Models (SpeechLLMs), similar to conventional speech emotion recognition (SER) systems, still treat emotion understanding as a simple classification problem. This provides limited inte…

Cited by 0SourcecodeScholar
2026

JAEGER: Joint 3D Audio-Visual Grounding and Reasoning in Simulated Physical Environments

ICML 2026poster

Current audio-visual large language models (AV-LLMs) are predominantly restricted to 2D perception, relying on RGB video and monaural audio. This design choice introduces a fundamental dimensionality mismatch that precludes reliable source localization and spatial reasoning in complex 3D environment…

Cited by 0SourceScholar
2026

MULTI-CHANNEL SPEECH ENHANCEMENT FOR COCKTAIL PARTY SPEECH EMOTION RECOGNITION

ICASSP 2026poster

This paper highlights the critical importance of multi-channel speech enhancement (MCSE) for speech emotion recognition (ER) in cocktail party scenarios. A multi-channel speech dereverberation and separation front-end integrating DNN-WPE and mask-based MVDR is used to extract the target speaker's sp…

Cited by 0SourcePDFScholar
2025

Effective and Efficient Mixed Precision Quantization of Speech Foundation Models

ICASSP 2025accepted

This paper presents a novel mixed-precision quantization approach for speech foundation models that tightly integrates mixed-precision learning and quantized model parameter estimation into one single model compression stage. Experiments conducted on LibriSpeech dataset with fine-tuned wav2vec2.0-ba…

Cited by 7SourceScholar
2025

Phone-purity Guided Discrete Tokens for Dysarthric Speech Recognition

ICASSP 2025accepted

Discrete tokens provide compact and domain-adaptable representations of speech features. However, their application to disordered speech, characterized by articulation imprecision and significant mismatch with normal voice, remains unexplored. To this end, this paper proposes novel phone-purity guid…

Cited by 0SourceScholar