← Search

Boyuan Cao

3 accepted papers

2026

Hierarchical Codec Diffusion for Video-to-Speech Generation

CVPR 2026

Video-to-Speech (VTS) generation aims to synthesize speech from a silent video without auditory signals, and holds substantial promise for applications such as film dubbing and voice restoration for individuals with aphonia. However, existing VTS methods disregard the hierarchical nature of speech,

Cited by 0SourcecodeScholar
2025

RepLDM: Reprogramming Pretrained Latent Diffusion Models for High-Quality, High-Efficiency, High-Resolution Image Generation

NeurIPS 2025spotlight

While latent diffusion models (LDMs), such as Stable Diffusion, are designed for high-resolution image generation, they often struggle with significant structural distortions when generating images at resolutions higher than their training one. Instead of relying on extensive retraining, a more res…

Cited by 0SourcecodeScholar