← Search

Chenjie Cao

23 accepted papers

2026

EarthCrafter: Scalable 3D Earth Generation via Dual-Sparse Latent Diffusion

AAAI 2026technical

Despite the remarkable developments achieved by recent 3D generation works, scaling these methods to geographic extents, such as modeling thousands of square kilometers of Earth’s surface, remains an open challenge. We address this through a dual innovation in data infrastructure and model architect

Cited by 0SourcePDFScholar
2026

RealisMotion: Decomposed Human Motion Control and Video Generation in the World Space

ICML 2026poster

Generating human videos with realistic and controllable motions is a challenging task. While existing methods can generate visually compelling videos, they lack separate control over four key video elements: foreground subject, background video, human trajectory, and action patterns. In this paper, …

Cited by 0SourceScholar
2026

WorldCompass: Reinforcement Learning for Long-Horizon World Models

ICML 2026poster

This work presents WorldCompass, a novel Reinforcement Learning (RL) post-training framework for the long-horizon, interactive video-based world models, enabling them to explore the world more accurately and consistently based on interaction signals. To effectively "steer" the world model's explorat…

Cited by 0SourceScholar
2026

WorldStereo: Bridging Camera-Guided Video Generation and Scene Reconstruction via 3D Geometric Memories

CVPR 2026

Recent advances in foundational Video Diffusion Models (VDMs) have yielded significant progress. Yet, despite the remarkable visual quality of generated videos, reconstructing consistent 3D scenes from these outputs remains challenging, due to limited camera controllability and inconsistent generate

Cited by 0SourcecodeScholar
2025

LiON-LoRA: Rethinking LoRA Fusion to Unify Controllable Spatial and Temporal Generation for Video Diffusion

ICCV 2025poster

Video Diffusion Models (VDMs) have demonstrated remarkable capabilities in synthesizing realistic videos by learning from large-scale data. Although vanilla Low-Rank Adaptation (LoRA) can learn specific spatial or temporal movement to driven VDMs with constrained data, achieving precise control over…

Cited by 0SourcePDFScholar
2025

MVGenMaster: Scaling Multi-View Generation from Any Image via 3D Priors Enhanced Diffusion Model

CVPR 2025poster

We introduce MVGenMaster, a multi-view diffusion model enhanced with 3D priors to address versatile Novel View Synthesis (NVS) tasks. MVGenMaster leverages 3D priors that are warped using metric depth and camera poses, significantly enhancing both generalization and 3D consistency in NVS. Our model…

2025

Towards Enhanced Image Inpainting: Mitigating Unwanted Object Insertion and Preserving Color Consistency

CVPR 2025highlight

Recent advances in image inpainting increasingly use generative models to handle large irregular masks. However, these models can create unrealistic inpainted images due to two main issues: (1) Unwanted object insertion: Even with unmasked areas as context, generative models may still generate arbit…

2024

Animate3D: Animating Any 3D Model with Multi-view Video Diffusion

NeurIPS 2024poster

Recent advances in 4D generation mainly focus on generating 4D content by distilling pre-trained text or single-view image conditioned models. It is inconvenient for them to take advantage of various off-the-shelf 3D assets with multi-view attributes, and their results suffer from spatiotemporal inc…

Cited by 13SourcePDFScholar
2024

Improving Neural Surface Reconstruction with Feature Priors from Multi-View Images

ECCV 2024poster

"Recent advancements in Neural Surface Reconstruction (NSR) have significantly improved multi-view reconstruction when coupled with volume rendering. However, relying solely on photometric consistency in image space falls short of addressing complexities posed by real-world data, including occlusion…

2024

LeftRefill: Filling Right Canvas based on Left Reference through Generalized Text-to-Image Diffusion Model

CVPR 2024poster

This paper introduces LeftRefill an innovative approach to efficiently harness large Text-to-Image (T2I) diffusion models for reference-guided image synthesis. As the name implies LeftRefill horizontally stitches reference and target views together as a whole input. The reference image occupies the…

2024

MVInpainter: Learning Multi-View Consistent Inpainting to Bridge 2D and 3D Editing

NeurIPS 2024poster

Novel View Synthesis (NVS) and 3D generation have recently achieved prominent improvements. However, these works mainly focus on confined categories or synthetic 3D assets, which are discouraged from generalizing to challenging in-the-wild scenes and fail to be employed with 2D synthesis directly. M…

2024

MVSFormer++: Revealing the Devil in Transformer's Details for Multi-View Stereo

ICLR 2024poster

Recent advancements in learning-based Multi-View Stereo (MVS) methods have prominently featured transformer-based models with attention mechanisms. However, existing approaches have not thoroughly investigated the profound influence of transformers on different MVS modules, resulting in limited dept…

2024

SC4D: Sparse-Controlled Video-to-4D Generation and Motion Transfer

ECCV 2024poster

"Recent advances in 2D/3D generative models enable the generation of dynamic 3D objects from a single-view video. Existing approaches utilize score distillation sampling to form the dynamic scene as dynamic NeRF or dense 3D Gaussians. However, these methods struggle to strike a balance among referen…

2024

VCD-Texture: Variance Alignment based 3D-2D Co-Denoising for Text-Guided Texturing

ECCV 2024poster

"Recent research on texture synthesis for 3D shapes benefits a lot from dramatically developed 2D text-to-image diffusion models, including inpainting-based and optimization-based approaches. However, these methods ignore the modal gap between the 2D diffusion model and 3D objects, which primarily r…

2023

Improving Transformer-based Image Matching by Cascaded Capturing Spatially Informative Keypoints

ICCV 2023poster

Learning robust local image feature matching is a fundamental low-level vision task, which has been widely explored in the past few years. Recently, detector-free local feature matchers based on transformers have shown promising results, which largely outperform pure Convolutional Neural Network (CN…

Cited by 11PDFcodeScholar
2023

Rethinking Optical Flow From Geometric Matching Consistent Perspective

CVPR 2023poster

Optical flow estimation is a challenging problem remaining unsolved. Recent deep learning based optical flow models have achieved considerable success. However, these models often train networks from the scratch on standard optical flow data, which restricts their ability to robustly and geometrical…

2022

High-Fidelity Portrait Editing Via Exploring Differentiable Guided Sketches from the Latent Space

ICASSP 2022accepted

This paper studies the task of sketch-guided high-fidelity portrait editing. Advanced unconditional generators, such as StyleGAN, can generate a high-quality portrait image with great diversity. In previous researches, StyleGAN has successfully been utilized for color-guided image editing through la…

Cited by 0SourceScholar
2022

Incremental Transformer Structure Enhanced Image Inpainting With Masking Positional Encoding

CVPR 2022poster

Image inpainting has made significant advances in recent years. However, it is still challenging to recover corrupted images with both vivid textures and reasonable structures. Some specific methods can only tackle regular textures while losing holistic structures due to the limited receptive fields…

Cited by 201PDFcodeScholar
2021

The Image Local Autoregressive Transformer

NeurIPS 2021poster

Recently, AutoRegressive (AR) models for the whole image generation empowered by transformers have achieved comparable or even better performance compared to Generative Adversarial Networks (GANs). Unfortunately, directly applying such AR models to edit/change local image regions, may suffer from th…

Cited by 13SourcePDFScholar
2020

CLUE: A Chinese Language Understanding Evaluation Benchmark

COLING 2020main

The advent of natural language understanding (NLU) benchmarks for English, such as GLUE and SuperGLUE allows new NLU models to be evaluated across a diverse set of tasks. These comprehensive benchmarks have facilitated a broad range of research and applications in natural language processing (NLP).…