← Search

Angtian Wang

23 accepted papers

2026

HECTOR: Hybrid Editable Compositional Object References for Video Generation

ICML 2026poster

Real-world videos naturally portray complex interactions among distinct physical objects, effectively forming dynamic compositions of visual elements. However, most current video generation models synthesize scenes holistically and therefore lack mechanisms for explicit compositional manipulation. T…

Cited by 0SourceScholar
2026

MAGREF: Masked Guidance for Any-Reference Video Generation with Subject Disentanglement

ICLR 2026poster

We tackle the task of any-reference video generation, which aims to synthesize videos conditioned on arbitrary types and combinations of reference subjects, together with textual prompts. This task faces persistent challenges, including identity inconsistency, entanglement among multiple reference s…

Cited by 0SourcecodeScholar
2026

TGT: Text-Grounded Trajectories for Locally Controlled Video Generation

CVPR 2026

Text-to-video generation has advanced rapidly in visual fidelity, whereas standard methods still have limited ability to control the subject composition of generated scenes. Prior work shows that adding localized text control signals, such as bounding boxes or segmentation masks, can help. However,

Cited by 0SourceScholar
2026

VIVA: VLM-Guided Instruction-Based Video Editing with Reward Optimization

CVPR 2026

Instruction-based video editing aims to modify an input video according to a natural-language instruction while preserving content fidelity and temporal coherence. However, existing diffusion-based approaches are often trained on paired data of simple editing operations, which fundamentally limits t

Cited by 0SourcecodeScholar
2025

Adventurer: Optimizing Vision Mamba Architecture Designs for Efficiency

CVPR 2025poster

In this work, we introduce the Adventurer series models where we treat images as sequences of patch tokens and employ uni-directional language models to learn visual representations. This modeling paradigm allows us to process images in a recurrent formulation with linear complexity relative to the…

Cited by 0SourcePDFScholar
2025

Compositional 4D Dynamic Scenes Understanding with Physics Priors for Video Question Answering

ICLR 2025poster

For vision-language models (VLMs), understanding the dynamic properties of objects and their interactions in 3D scenes from videos is crucial for effective reasoning about high-level temporal and action semantics. Although humans are adept at understanding these properties by constructing 3D and tem…

2025

PartInstruct: Part-level Instruction Following for Fine-grained Robot Manipulation

RSS 2025poster

Fine-grained robot manipulation, such as lifting and rotating a bottle to display the label on the cap, requires robust reasoning about object parts and their relationships with intended tasks. Despite recent advances in training general-purpose robot manipulation policies guided by language instruc…

Cited by 0PDFScholar
2025

WorldWeaver: Generating Long-Horizon Video Worlds via Rich Perception

NeurIPS 2025poster

Generative video modeling has made significant strides, yet ensuring structural and temporal consistency over long sequences remains a challenge. Current methods predominantly rely on RGB signals, leading to accumulated errors in object structure and motion over extended durations. To address these…

Cited by 0SourceScholar
2024

Generating Images with 3D Annotations Using Diffusion Models

ICLR 2024spotlight

Diffusion models have emerged as a powerful generative method, capable of producing stunning photo-realistic images from natural language descriptions. However, these models lack explicit control over the 3D structure in the generated images. Consequently, this hinders our ability to obtain detailed…

Cited by 6SourcePDFScholar
2024

HISR: Hybrid Implicit Surface Representation for Photorealistic 3D Human Reconstruction

AAAI 2024technical

Neural reconstruction and rendering strategies have demonstrated state-of-the-art performances due, in part, to their ability to preserve high level shape details. Existing approaches, however, either represent objects as implicit surface functions or neural volumes and still struggle to recover sha…

Cited by 3SourcePDFScholar
2024

NOVUM: Neural Object Volumes for Robust Object Classification

ECCV 2024poster

"Discriminative models for object classification typically learn image-based representations that do not capture the compositional and 3D nature of objects. In this work, we show that explicitly integrating 3D compositional object representations into deep networks for image classification leads to…

2024

Radiative Gaussian Splatting for Efficient X-ray Novel View Synthesis

ECCV 2024poster

"X-ray is widely applied for transmission imaging due to its stronger penetration than natural light. When rendering novel view X-ray projections, existing methods mainly based on NeRF suffer from long training time and slow inference speed. In this paper, we propose a 3D Gaussian splatting-based me…

2024

Semantic Flow: Learning Semantic Fields of Dynamic Scenes from Monocular Videos

ICLR 2024poster

In this work, we pioneer Semantic Flow, a neural semantic representation of dynamic scenes from monocular videos. In contrast to previous NeRF methods that reconstruct dynamic scenes from the colors and volume densities of individual points, Semantic Flow learns semantics from continuous flows that…

Cited by 5SourcePDFScholar
2024

Structure-Aware Sparse-View X-ray 3D Reconstruction

CVPR 2024poster

X-ray known for its ability to reveal internal structures of objects is expected to provide richer information for 3D reconstruction than visible light. Yet existing NeRF algorithms overlook this nature of X-ray leading to their limitations in capturing structural contents of imaged objects. In this…

2024

iNeMo: Incremental Neural Mesh Models for Robust Class-Incremental Learning

ECCV 2024poster

"Different from human nature, it is still common practice today for vision tasks to train deep learning models only initially and on fixed datasets. A variety of approaches have recently addressed handling continual data streams. However, extending these methods to manage out-of-distribution (OOD) s…

2023

3D-Aware Neural Body Fitting for Occlusion Robust 3D Human Pose Estimation

ICCV 2023poster

Regression-based methods for 3D human pose estimation directly predict the 3D pose parameters from a 2D image using deep networks. While achieving state-of-the-art performance on standard benchmarks, their performance degrades under occlusion. In contrast, optimization-based methods fit a parametric…

Cited by 42PDFcodeScholar
2023

VoGE: A Differentiable Volume Renderer using Gaussian Ellipsoids for Analysis-by-Synthesis

ICLR 2023poster

Differentiable rendering allows the application of computer graphics on vision tasks, e.g. object pose and shape fitting, via analysis-by-synthesis, where gradients at occluded regions are important when inverting the rendering process.To obtain those gradients, state-of-the-art (SoTA) differentiabl…

2022

OOD-CV: A Benchmark for Robustness to Out-of-Distribution Shifts of Individual Nuisances in Natural Images

ECCV 2022poster

"Enhancing the robustness of vision algorithms in real-world scenarios is challenging. One reason is that existing robustness benchmarks are limited, as they either rely on synthetic data or ignore the effects of individual nuisance factors. We introduce ROBIN, a benchmark dataset that includes out-…

2022

Robust Category-Level 6D Pose Estimation with Coarse-to-Fine Rendering of Neural Features

ECCV 2022poster

"We consider the problem of category-level 6D pose estimation from a single RGB image. Our approach represents an object category as a cuboid mesh and learns a generative model of the neural feature activations at each mesh vertex to perform pose estimation through differentiable rendering. A common…

2021

NeMo: Neural Mesh Models of Contrastive Features for Robust 3D Pose Estimation

ICLR 2021poster

3D pose estimation is a challenging but important task in computer vision. In this work, we show that standard deep learning approaches to 3D pose estimation are not robust to partial occlusion. Inspired by the robustness of generative vision models to partial occlusion, we propose to integrate deep…

2021

Neural View Synthesis and Matching for Semi-Supervised Few-Shot Learning of 3D Pose

NeurIPS 2021poster

We study the problem of learning to estimate the 3D object pose from a few labelled examples and a collection of unlabelled data. Our main contribution is a learning framework, neural view synthesis and matching, that can transfer the 3D pose annotation from the labelled to unlabelled images reliabl…

2020

Robust Object Detection Under Occlusion With Context-Aware CompositionalNets

CVPR 2020poster

Detecting partially occluded objects is a difficult task. Our experimental results show that deep learning approaches, such as Faster R-CNN, are not robust at object detection under occlusion. Compositional convolutional neural networks (CompositionalNets) have been shown to be robust at classifying…

Cited by 163PDFScholar
2018

Weakly Supervised Region Proposal Network and Object Detection

ECCV 2018poster

The Convolutional Neural Network (CNN) based region proposal generation method (i.e. region proposal network), trained using bounding box annotations, is an essential component in modern fully supervised object detectors. However, Weakly Supervised Object Detection (WSOD) has not benefited from CNN-…

Cited by 247SourcePDFScholar