← Search

Weijian Cao

13 accepted papers

2026

Dual Latent Memory for Visual Multi-agent System

ICML 2026poster

While Visual Multi-Agent Systems (VMAS) promise to enhance comprehensive abilities through inter-agent collaboration, empirical evidence reveals a counter-intuitive "scaling wall": increasing agent turns often degrades performance while exponentially inflating token costs. We attribute this failure …

Cited by 0SourceScholar
2026

FeRA: Frequency-Energy Constrained Routing for Effective Diffusion Adaptation Fine-Tuning

ICML 2026poster

Diffusion models have achieved remarkable success in generative modeling, yet how to effectively adapting large pretrained models to new tasks remains challenging. We revisit the reconstruction behavior of diffusion models during denoising to unveil the underlying frequency–energy mechanism governin…

Cited by 0SourceScholar
2026

Reasoning to Edit: Hypothetical Instruction-Based Image Editing with Visual Reasoning

ICML 2026poster

Instruction-based image editing (IIE) has advanced rapidly with the success of diffusion models. However, existing efforts primarily focus on simple and explicit instructions to execute editing operations such as adding, deleting, moving, or swapping objects. They struggle to handle more complex imp…

Cited by 0SourceScholar
2026

Soul: Breathe Life into Digital Human for High-fidelity Long-term Multimodal Animation

CVPR 2026

We propose a multimodal-driven framework for high-fidelity long-term digital human animation termed Soul, which generates semantically coherent videos from a single-frame portrait image, text prompts, and audio, achieving precise lip synchronization, vivid facial expressions, and robust identity pre

Cited by 0SourceScholar
2026

SwiftVideo: A Unified Framework for Few-Step Video Generation Through Trajectory-Distribution Alignment

AAAI 2026technical

Diffusion-based or flow-based models have achieved significant progress in video synthesis but require multiple iterative sampling steps, which incurs substantial computational overhead. While many distillation methods that are solely based on trajectory-preserving or distribution-matching have been

Cited by 0SourcePDFScholar
2026

TGPO: Efficient Policy Optimization through Sequence Anchor and Information Gating

ICML 2026poster

Reinforcement learning from verifiable rewards (RLVR) has become an important paradigm for enhancing the reasoning capabilities of large language models, while it also involves a persistent tradeoff between optimization stability and learning efficiency. Token-level importance weighting supports fin…

Cited by 0SourceScholar
2026

Transform Trained Transformer for Accelerating Native 4K Video Generation

ICML 2026poster

Native 4K (2176$\times$3840) video generation remains a critical challenge due to the quadratic computational explosion of full-attention as spatiotemporal resolution increases, making it difficult for models to strike a balance between efficiency and quality. This paper proposes a novel Transformer…

Cited by 0SourceScholar
2026

Transform Trained Transformer for Accelerating Native 4K Video Generation

ICML 2026poster

Native 4K (2176$\times$3840) video generation remains a critical challenge due to the quadratic computational explosion of full-attention as spatiotemporal resolution increases, making it difficult for models to strike a balance between efficiency and quality. This paper proposes a novel Transformer…

Cited by 0SourceScholar
2025

ID-Sculpt: ID-aware 3D Head Generation from Single In-the-wild Portrait Image

AAAI 2025technical

While recent works have achieved great success on one-shot 3D common object generation, high quality and fidelity 3D head generation from a single image remains a great challenge. Previous text-based methods for generating 3D heads were limited by text descriptions and image-based methods struggled…

Cited by 0SourcePDFScholar
2024

FreeMotion: A Unified Framework for Number-free Text-to-Motion Synthesis

ECCV 2024poster

"Text-to-motion synthesis is a crucial task in computer vision. Existing methods are limited in their universality, as they are tailored for single-person or two-person scenarios and can not be applied to generate motions for more individuals. To achieve the number-free motion synthesis, this paper…

Cited by 19SourcePDFScholar
2024

TexDreamer: Towards Zero-Shot High-Fidelity 3D Human Texture Generation

ECCV 2024oral

"Texturing 3D humans with semantic UV maps remains a challenge due to the difficulty of acquiring reasonably unfolded UV. Despite recent text-to-3D advancements in supervising multi-view renderings using large text-to-image (T2I) models, issues persist with generation speed, text consistency, and te…

Cited by 8SourcePDFScholar
2023

Learning Neural Proto-Face Field for Disentangled 3D Face Modeling in the Wild

CVPR 2023poster

Generative models show good potential for recovering 3D faces beyond limited shape assumptions. While plausible details and resolutions are achieved, these models easily fail under extreme conditions of pose, shadow or appearance, due to the entangled fitting or lack of multi-view priors. To address…

Cited by 6SourcePDFScholar
2022

Physically-Guided Disentangled Implicit Rendering for 3D Face Modeling

CVPR 2022poster

This paper presents a novel Physically-guided Disentangled Implicit Rendering (PhyDIR) framework for high-fidelity 3D face modeling. The motivation comes from two observations: widely-used graphics renderers yield excessive approximations against photo-realistic imaging, while neural rendering metho…

Cited by 8PDFScholar