← Search

Shuaiqi Chen

3 accepted papers

2025

DecoupledSynth: Enhancing Zero-Shot Text-to-Speech Via Factors Decoupling

ICASSP 2025accepted

Studies of speech representation enhance zero-shot Text-to-Speech by mapping text to intermediate representations before generating speech. However, using representations often struggles to balance linguistic, para-linguistic, and non-linguistic information in speech during the synthesis phase. Addi…

Cited by 0SourceScholar
2024

Draw Step by Step: Reconstructing CAD Construction Sequences from Point Clouds via Multimodal Diffusion.

CVPR 2024poster

Reconstructing CAD construction sequences from raw 3D geometry serves as an interface between real-world objects and digital designs. In this paper we propose CAD-Diffuser a multimodal diffusion scheme aiming at integrating top-down design paradigm into generative reconstruction. In particular we un…

Cited by 10SourcePDFScholar
2023

DWFormer: Dynamic Window Transformer for Speech Emotion Recognition

ICASSP 2023accepted

Speech emotion recognition is crucial to human-computer interaction. The temporal regions that represent different emotions scatter in different parts of the speech locally. Moreover, the temporal scales of important information may vary over a large range within and across speech segments. Although…

Cited by 0SourceScholar