← Search

Sizhe Shan

3 accepted papers

2026

TMD-Bench: A Multi-Level Evaluation Paradigm for Music–Dance Co-Generation

ICML 2026poster

Unified audio--visual generation is rapidly gaining industrial and creative relevance, enabling applications in virtual production and interactive media. However, when moving from general audio--video synthesis to music–dance co-generation, the task becomes substantially harder: musical rhythm, phra…

Cited by 0SourceScholar
2025

Do Less and Achieve More: Free Condition Video Outpainting with Diffusion Model

ICASSP 2025accepted

Video outpainting aims to extend the content of a video beyond its original spatial boundaries. Existing methods tend to condition the generation process on a single frame or caption, failing to address the challenge in long videos with multiple video clips. To address this, we extend the diffusion-…

Cited by 0SourceScholar
2025

ProsodyFlow: High-fidelity Text-to-Speech through Conditional Flow Matching and Prosody Modeling with Large Speech Language Models

COLING 2025main

Text-to-speech (TTS) has seen significant advancements in high-quality, expressive speech synthesis. However, achieving diverse and natural prosody in synthesized speech remains challenging. In this paper, we propose ProsodyFlow, an end-to-end TTS model that integrates large self-supervised speech m…

Cited by 0SourcePDFScholar