← Search

Chaoyang Wang

35 accepted papers

2026

EasyV2V: A High-quality Instruction-based Video Editing Framework

CVPR 2026

While image editing has advanced rapidly, video editing remains less explored, facing challenges in consistency, control, and generalization.We study the design space of data, architecture, and control, and introduce EasyV2V, a simple and effective framework for instruction-based video editing. On t

Cited by 4SourcecodeScholar
2026

EgoEdit: Dataset, Real-Time Streaming Model, and Benchmark for Egocentric Video Editing

CVPR 2026

We study instruction-guided editing of egocentric videos for interactive AR applications. While recent AI video editors perform well on third-person footage, egocentric views present unique challenges -- including rapid egomotion, and frequent hand-object interactions -- that create a significant do

Cited by 0SourcecodeScholar
2026

RigMo: Unifying Rig and Motion Learning for Generative Animation

CVPR 2026

Despite significant progress in 4D generation, rig and motion--the core structural and dynamic components of animation--are typically modeled as separate problems. Existing pipelines rely on ground-truth skeletons and skinning weights for motion generation and treat auto-rigging as an independent pr

Cited by 0SourceScholar
2026

ShapeGen4D: Towards High Quality 4D Shape Generation from Videos

ICLR 2026poster

Video-conditioned 4D shape generation aims to recover time-varying 3D geometry and view-consistent appearance directly from an input video. In this work, we introduce a native video-to-4D shape generation framework that synthesizes a single dynamic 3D representation end-to-end from the video. Our…

Cited by 0SourceScholar
2025

3DitScene: Editing Any Scene via Language-guided Disentangled Gaussian Splatting

ICLR 2025poster

Scene image editing is crucial for entertainment, photography, and advertising design. Existing methods solely focus on either 2D individual object or 3D global scene editing. This results in a lack of a unified approach to effectively control and manipulate scenes at the 3D level with different lev…

Cited by 4SourcePDFScholar
2025

4Real-Video: Learning Generalizable Photo-Realistic 4D Video Diffusion

CVPR 2025highlight

We propose 4Real-Video, a novel framework for generating 4D videos, organized as a grid of video frames with both time and viewpoint axes. In this grid, each row contains frames sharing the same timestep, while each column contains frames from the same viewpoint. One stream performs viewpoint updat…

Cited by 2SourcePDFScholar
2025

Conditional Panoramic Image Generation via Masked Autoregressive Modeling

NeurIPS 2025poster

Recent progress in panoramic image generation has underscored two critical limitations in existing approaches. First, most methods are built upon diffusion models, which are inherently ill-suited for equirectangular projection (ERP) panoramas due to the violation of the identically and independently…

Cited by 0SourceScholar
2025

DELTA: DENSE EFFICIENT LONG-RANGE 3D TRACKING FOR ANY VIDEO

ICLR 2025poster

Tracking dense 3D motion from monocular videos remains challenging, particularly when aiming for pixel-level precision over long sequences. We introduce DELTA, a novel method that efficiently tracks every pixel in 3D space, enabling accurate motion estimation across entire videos. Our approach lever…

Cited by 4SourcePDFScholar
2025

Explore In-Context Segmentation via Latent Diffusion Models

AAAI 2025technical

In-context segmentation has drawn increasing attention with the advent of vision foundation models. Its goal is to segment objects using given reference images. Most existing approaches adopt metric learning or masked image modeling to build the correlation between visual prompts and input image que…

Cited by 10SourcePDFScholar
2025

Fused View-Time Attention and Feedforward Reconstruction for 4D Scene Generation

NeurIPS 2025poster

We propose the first framework capable of computing a 4D spatio-temporal grid of video frames and 3D Gaussian particles for each time step using a feed-forward architecture. Our architecture has two main components, a 4D video model and a 4D reconstruction model. In the first part, we analyze curren…

Cited by 0SourceScholar
2025

GTR: Improving Large 3D Reconstruction Models through Geometry and Texture Refinement

ICLR 2025poster

We propose a novel approach for 3D mesh reconstruction from multi-view images. We improve upon the large reconstruction model LRM that use a transformer-based triplane generator and a Neural Radiance Field (NeRF) model trained on multi-view images. We introduce three key components to significantly…

Cited by 3SourcePDFScholar
2025

Lightweight Predictive 3D Gaussian Splats

ICLR 2025poster

Recent approaches representing 3D objects and scenes using Gaussian splats show increased rendering speed across a variety of platforms and devices. While rendering such representations is indeed extremely efficient, storing and transmitting them is often prohibitively expensive. To represent large-…

2025

OracleFusion: Assisting the Decipherment of Oracle Bone Script with Structurally Constrained Semantic Typography

ICCV 2025poster

As one of the earliest ancient languages, Oracle Bone Script (**OBS**) encapsulates the cultural records and intellectual expressions of ancient civilizations. Despite the discovery of approximately 4,500 OBS characters, only about 1,600 have been deciphered. The remaining undeciphered ones, with th…

2025

PrEditor3D: Fast and Precise 3D Shape Editing

CVPR 2025poster

We propose a training-free approach to 3D editing that enables the editing of a single shape and the reconstruction of a mesh within a few minutes. Leveraging 4-view images, user-guided text prompts, and rough 2D masks, our method produces an edited 3D mesh that aligns with the prompt. For this, our…

Cited by 3SourcePDFScholar
2025

T2Bs: Text-to-Character Blendshapes via Video Generation

ICCV 2025poster

We present T2Bs, a framework for generating high-quality, animatable character head morphable models from text by combining static text-to-3D generation with video diffusion. Text-to-3D models produce detailed static geometry but lack motion synthesis, while video diffusion models generate motion wi…

Cited by 0SourcePDFScholar
2025

VD3D: Taming Large Video Diffusion Transformers for 3D Camera Control

ICLR 2025poster

Modern text-to-video synthesis models demonstrate coherent, photorealistic generation of complex videos from a text description. However, most existing models lack fine-grained control over camera movement, which is critical for downstream applications related to content creation, visual effects, an…

Cited by 38SourcePDFScholar
2024

4Real: Towards Photorealistic 4D Scene Generation via Video Diffusion Models

NeurIPS 2024poster

Existing dynamic scene generation methods mostly rely on distilling knowledge from pre-trained 3D generative models, which are typically fine-tuned on synthetic object datasets. As a result, the generated scenes are often object-centric and lack photorealism. To address these limitations, we introd…

Cited by 26SourcePDFScholar
2024

LGMRec: Local and Global Graph Learning for Multimodal Recommendation

AAAI 2024technical

The multimodal recommendation has gradually become the infrastructure of online media platforms, enabling them to provide personalized service to users through a joint modeling of user historical behaviors (e.g., purchases, clicks) and item various modalities (e.g., visual and textual). The majority…

2024

SemFlow: Binding Semantic Segmentation and Image Synthesis via Rectified Flow

NeurIPS 2024poster

Semantic segmentation and semantic image synthesis are two representative tasks in visual perception and generation. While existing methods consider them as two distinct tasks, we propose a unified framework (SemFlow) and model them as a pair of reverse problems. Specifically, motivated by rectified…

2024

Towards Text-guided 3D Scene Composition

CVPR 2024poster

We are witnessing significant breakthroughs in the technology for generating 3D objects from text. Existing approaches either leverage large text-to-image models to optimize a 3D representation or train 3D generators on object-centric datasets. Generating entire scenes however remains very challengi…

2023

Autodecoding Latent 3D Diffusion Models

NeurIPS 2023poster

Diffusion-based methods have shown impressive visual results in the text-to-image domain. They first learn a latent space using an autoencoder, then run a denoising process on the bottleneck to generate new samples. However, learning an autoencoder requires substantial data in the target domain. Suc…

2023

Controllable Chest X-Ray Report Generation from Longitudinal Representations

EMNLP 2023long findings

Radiology reports are detailed text descriptions of the content of medical scans. Each report describes the presence/absence and location of relevant clinical findings, commonly including comparison with prior exams of the same patient to describe how they evolved. Radiology reporting is a time-cons…

Cited by 0SourceScholar
2023

LightSpeed: Light and Fast Neural Light Fields on Mobile Devices

NeurIPS 2023poster

Real-time novel-view image synthesis on mobile devices is prohibitive due to the limited computational power and storage. Using volumetric rendering methods, such as NeRF and its derivatives, on mobile devices is not suitable due to the high computational cost of volumetric rendering. On the other h…

Cited by 11SourcePDFScholar
2023

Reconstructing Animatable Categories From Videos

CVPR 2023poster

Building animatable 3D models is challenging due to the need for 3D scans, laborious registration, and manual rigging. Recently, differentiable rendering provides a pathway to obtain high-quality 3D models from monocular videos, but these are limited to rigid categories or single instances. We prese…

2022

A polynomial time approximation scheme for the scheduling problem in the AGV system

IROS 2022poster

Logistics warehouses face the challenge of fulfilling large bulk pick orders limit in a given time, as the information of logistics orders is different and timeliness. Therefore, in automated warehouses, it is imperative to improve the efficiency and intelligence of order picking by robotic systems.…

Cited by 0SourceScholar
2022

MBW: Multi-view Bootstrapping in the Wild

NeurIPS 2022accept

Labeling articulated objects in unconstrained settings has a wide variety of applications including entertainment, neuroscience, psychology, ethology, and many fields of medicine. Large offline labeled datasets do not exist for all but the most common articulated object categories (e.g., humans). Ha…

Cited by 2SourcePDFScholar
2020

SDF-SRN: Learning Signed Distance 3D Object Reconstruction from Static Images

NeurIPS 2020poster

Dense 3D object reconstruction from a single image has recently witnessed remarkable advances, but supervising neural networks with ground-truth 3D shapes is impractical due to the laborious process of creating paired image-shape datasets. Recent efforts have turned to learning 3D reconstruction wit…

2018

Learning Depth From Monocular Videos Using Direct Methods

CVPR 2018poster

The ability to predict depth from a single image - using recent advances in CNNs - is of increasing interest to the vision community. Unsupervised strategies to learning are particularly appealing as they can utilize much larger and varied monocular video datasets during learning without the need fo…

2017

Rethinking Reprojection: Closing the Loop for Pose-Aware Shape Reconstruction From a Single Image

ICCV 2017spotlight

An emerging problem in computer vision is the reconstruction of 3D shape and pose of an object from a single image. Hitherto, the problem has been addressed through the application of canonical deep learning methods to regress from the image directly to the 3D shape and pose labels. These approaches…

Cited by 121PDFScholar
2015

Object Proposal by Multi-Branch Hierarchical Segmentation

CVPR 2015poster

Hierarchical segmentation based object proposal methods have become an important step in modern object detection paradigm. However, standard single-way hierarchical methods are fundamentally flawed in that the errors in early steps cannot be corrected and accumulate. In this work, we propose a novel…

Cited by 45SourcePDFScholar