← Search

Yushen Chen

4 accepted papers

2026

AUV: TEACHING AUDIO UNIVERSAL VECTOR QUANTIZATION WITH SINGLE NESTED CODEBOOK

ICASSP 2026poster

We propose AUV, a unified neural audio codec with a single codebook, which enables a favourable reconstruction of speech and further extends to general audio, including vocal, music, and sound. AUV is capable of tackling any 16 kHz mixed-domain audio segment at bit rates around 700 bps. To accomplis…

Cited by 0SourcePDFScholar
2026

Cross-Lingual F5-TTS: Towards Language-Agnostic Voice Cloning and Speech Synthesis

ICASSP 2026poster

Flow-matching-based text-to-speech (TTS) models have shown high-quality speech synthesis. However, most current flow-matching-based TTS models still rely on reference transcripts corresponding to the audio prompt for synthesis. This dependency prevents cross-lingual voice cloning when audio prompt t…

Cited by 0SourcePDFScholar
2025

F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching

ACL 2025long

This paper introduces F5-TTS, a fully non-autoregressive text-to-speech system based on flow matching with Diffusion Transformer (DiT). Without requiring complex designs such as duration model, text encoder, and phoneme alignment, the text input is simply padded with filler tokens to the same length…

2023

An Anthropomorphic Robotic Hand With a Soft-Rigid Hybrid Structure and Positive- Negative Pneumatic Actuation

RA-L 2023

Anthropomorphic robotic hands are seeking to achieve key features such as multi-degree-of-freedom motion ability, bi-directional actuation, high adaptability, and sufficient stiffness. In this research, we propose a 10 active degrees-of-freedom anthropomorphic robotic hand with a soft-rigid hybrid s

Cited by 22SourceScholar