← Search

Zhuoliang Kang

7 accepted papers

2026

Active Intelligence in Video Avatars via Closed-loop World Modeling

CVPR 2026

Current video avatar generation methods excel at identity preservation and motion alignment but lack genuine agency--they cannot autonomously pursue long-term goals through adaptive environmental interaction. We address this by introducing L-IVA (Long-horizon Interactive Visual Avatar), a task and b

Cited by 0SourceScholar
2026

Infinite-World: Scaling Interactive World Models to 1000-Frame Horizons via Pose-Free Hierarchical Memory

ICML 2026poster

We propose **Infinite-World**, a robust interactive world model capable of maintaining coherent visual memory over **1000+ frames** in complex real-world environments. While existing world models can be efficiently optimized on synthetic data with perfect ground-truth, they lack an effective trainin…

Cited by 0SourceScholar
2026

LUVE : Latent-Cascaded Ultra-High-Resolution Video Generation with Dual Frequency Experts

ICML 2026poster

Recent advances in video diffusion models have significantly improved visual quality, yet ultra-high-resolution (UHR) video generation remains a formidable challenge due to the compounded difficulties of motion modeling, semantic planning, and detail synthesis. To address these limitations, we propo…

Cited by 0SourceScholar
2026

U-Mind: A Unified Framework for Real-Time Multimodal Interaction with Audiovisual Generation

CVPR 2026

Full-stack multimodal interaction in real-time is a central goal in building intelligent embodied agents capable of natural, dynamic communication. However, existing systems are either limited to unimodal generation or suffer from degraded reasoning and poor cross-modal alignment, preventing coheren

Cited by 0SourcecodeScholar
2026

WildActor: Unconstrained Identity-Preserving Video Generation

ICML 2026poster

Production-ready human video generation requires digital actors to maintain strictly consistent full-body identities across dynamic shots, viewpoints and motions, a setting that remains challenging for existing methods. Prior methods often suffer from face-centric behavior that neglects body-level c…

Cited by 0SourceScholar
2025

Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation

NeurIPS 2025poster

Audio-driven human animation methods, such as talking head and talking body generation, have made remarkable progress in generating synchronized facial movements and appealing visual quality videos. However, existing methods primarily focus on single human animation and struggle with multi-stream au…

Cited by 0SourcecodeScholar
2024

Rad-NeRF: Ray-decoupled Training of Neural Radiance Field

NeurIPS 2024poster

Although the neural radiance field (NeRF) exhibits high-fidelity visualization on the rendering task, it still suffers from rendering defects, especially in complex scenes. In this paper, we delve into the reason for the unsatisfactory performance and conjecture that it comes from interference in th…