← Search

Ke Cao

9 accepted papers

2026

Active Intelligence in Video Avatars via Closed-loop World Modeling

CVPR 2026

Current video avatar generation methods excel at identity preservation and motion alignment but lack genuine agency--they cannot autonomously pursue long-term goals through adaptive environmental interaction. We address this by introducing L-IVA (Long-horizon Interactive Visual Avatar), a task and b

Cited by 0SourceScholar
2026

Cross-Scale Pansharpening via ScaleFormer and the PanScale Benchmark

CVPR 2026

Pansharpening aims to generate high-resolution multi-spectral images by fusing the spatial detail of panchromatic images with the spectral richness of low-resolution MS data. However, most existing methods are evaluated under limited, low-resolution settings, limiting their generalization to real-wo

Cited by 0SourcecodeScholar
2026

InnoAds-Composer: Efficient Condition Composition for E-Commerce Poster Generation

CVPR 2026

E-commerce product poster generation aims to automatically synthesize a single image that effectively conveys product information by presenting a subject, text, and a designed style. Recent diffusion models with fine-grained and efficient controllability have advanced product poster synthesis, yet t

Cited by 0SourceScholar
2026

MoFu: Scale-Aware Modulation and Fourier Fusion for Multi-Subject Video Generation

AAAI 2026technical

Multi-subject video generation aims to synthesize videos from textual prompts and multiple reference images, ensuring that each subject preserves natural scale and visual fidelity. However, current methods face two challenges: scale inconsistency, where variations in subject size lead to unnatural g

Cited by 0SourcePDFScholar
2026

RelaCtrl: Relevance-Guided Efficient Control for Diffusion Transformers

AAAI 2026technical

The Diffusion Transformer plays a pivotal role in advancing text-to-image and text-to-video generation, owing primarily to its inherent scalability. However, existing controlled diffusion transformer methods incur significant parameter and computational overheads and suffer from inefficient resource

Cited by 0SourcePDFScholar
2026

Self-supervised Multiplex Consensus Mamba for General Image Fusion

AAAI 2026technical

Image fusion integrates complementary information from different modalities to generate high-quality fused images, thereby enhancing downstream tasks such as object detection and semantic segmentation. Unlike task-specific techniques that primarily focus on consolidating inter-modal information, gen

Cited by 0SourcePDFScholar
2025

FancyVideo: Towards Dynamic and Consistent Video Generation via Cross-frame Textual Guidance

IJCAI 2025

Synthesizing motion-rich and temporally consistent videos remains a challenge in artificial intelligence, especially when dealing with extended durations. Existing text-to-video (T2V) models commonly employ spatial cross-attention for text control, equivalently guiding different frame generations wi

Cited by 0SourcePDFScholar
2025

Lay2Story: Extending Diffusion Transformers for Layout-Togglable Story Generation

ICCV 2025poster

Storytelling tasks involving generating consistent subjects have gained significant attention recently. However, existing methods, whether training-free or training-based, continue to face challenges in maintaining subject consistency due to the lack of fine-grained guidance and inter-frame interact…

Cited by 0SourcePDFScholar
2025

WISA: World simulator assistant for physics-aware text-to-video generation

NeurIPS 2025spotlight

Recent advances in text-to-video (T2V) generation, exemplified by models such as Sora and Kling, have demonstrated strong potential for constructing world simulators. However, existing T2V models still struggle to understand abstract physical principles and to generate videos that faithfully obey ph…

Cited by 0SourcecodeScholar