Real-Time Generation of Streamable Talking Portrait Video with Reference-Guided Deep Compression VAEs
Video diffusion models have significantly advanced portrait video generation, yet their high computational demands limit their use in interactive applications. This work presents a framework for streamable talking portrait video generation conditioned on speech audio and reference images. Designed m