← Search

Sitong Cheng

3 accepted papers

2026

UniSS: Unified Expressive Speech-to-Speech Translation with Your Voice

ICLR 2026poster

The ultimate goal of expressive speech-to-speech translation (S2ST) is to accurately translate spoken content while preserving the speaker identity and emotional style. However, progress in this field is largely hindered by three key challenges: the scarcity of paired speech data that retains expres…

Cited by 0SourcecodeScholar
2025

Both Ears Wide Open: Towards Language-Driven Spatial Audio Generation

ICLR 2025spotlight

Recently, diffusion models have achieved great success in mono-channel audio generation. However, when it comes to stereo audio generation, the soundscapes often have a complex scene of multiple objects and directions. Controlling stereo audio with spatial contexts remains challenging due to high da…

Cited by 2SourcePDFScholar
2020

CN-Celeb: A Challenging Chinese Speaker Recognition Dataset

ICASSP 2020accepted

Recently, researchers set an ambitious goal of conducting speaker recognition in unconstrained conditions where the variations on ambient, channel and emotion could be arbitrary. However, most publicly available datasets are collected under constrained environments, i.e., with little noise and limit…

Cited by 271SourceScholar