← Search

YIYI CAI

2 accepted papers

2026

FloodDiffusion: Tailored Diffusion Forcing for Streaming Motion Generation

CVPR 2026

We present FloodDiffusion, a new framework for text-driven, streaming human motion generation. Given time-varying text prompts, FloodDiffusion generates text-aligned, seamless motion sequences with real-time latency.Unlike existing methods that rely on chunk-by-chunk or auto-regressive model with di

Cited by 0SourceScholar
2025

Shallow Flow Matching for Coarse-to-Fine Text-to-Speech Synthesis

NeurIPS 2025poster

We propose Shallow Flow Matching (SFM), a novel mechanism that enhances flow matching (FM)-based text-to-speech (TTS) models within a coarse-to-fine generation paradigm. Unlike conventional FM modules, which use the coarse representations from the weak generator as conditions, SFM constructs interme…

Cited by 0SourcecodeScholar