← Search

Xingang Pan

37 accepted papers

2026

4RC: 4D Reconstruction via Conditional Querying Anytime and Anywhere

ICML 2026poster

We present 4RC, a unified feed-forward framework for 4D reconstruction from monocular videos. Unlike existing methods that typically decouple motion from geometry or produce limited 4D attributes, such as sparse trajectories or two-view scene flow, 4RC learns a holistic 4D representation that jointl…

Cited by 0SourceScholar
2026

PI-Light: Physics-Inspired Diffusion for Full-Image Relighting

ICLR 2026poster

Full-image relighting remains a challenging problem due to the difficulty of collecting large-scale structured paired data, the difficulty of maintaining physical plausibility, and the limited generalizability imposed by data-driven priors. Existing attempts to bridge the synthetic-to-real gap for f…

Cited by 0SourcecodeScholar
2026

STream3R: Scalable Sequential 3D Reconstruction with Causal Transformer

ICLR 2026poster

We present STream3R, a novel approach to 3D reconstruction that reformulates pointmap prediction as a decoder-only Transformer problem. Existing state-of-the-art methods for multi-view reconstruction either depend on expensive global optimization or rely on simplistic memory mechanisms that scale po…

Cited by 0SourcecodeScholar
2026

Trainable Log-linear Sparse Attention for Efficient Diffusion Transformers

CVPR 2026

Diffusion Transformers (DiTs) set the state of the art in visual generation, yet their quadratic self-attention cost fundamentally limits scaling to long token sequences. Recent Top-K sparse attention approaches reduce the computation of DiTs by compressing tokens into block-wise representation and

Cited by 4SourceScholar
2025

3DEnhancer: Consistent Multi-View Diffusion for 3D Enhancement

CVPR 2025poster

Despite advances in neural rendering, due to the scarcity of high-quality 3D datasets and the inherent limitations of multi-view diffusion models, view synthesis and 3D model generation are restricted to low resolutions with suboptimal multi-view consistency. In this study, we present a novel 3D enh…

2025

Alias-Free Latent Diffusion Models: Improving Fractional Shift Equivariance of Diffusion Latent Space

CVPR 2025poster

Latent Diffusion Models (LDMs) are known to have an unstable generation process, where even small perturbations or shifts in the input noise can lead to significantly different outputs. This hinders their applicability in applications requiring consistent results. In this work, we redesign LDMs to e…

2025

Epona: Autoregressive Diffusion World Model for Autonomous Driving

ICCV 2025poster

Diffusion models have demonstrated exceptional visual quality in video generation, making them promising for autonomous driving world modeling. However, existing video diffusion-based world models struggle with flexible-length, long-horizon predictions and integrating trajectory planning. This is be…

2025

FreeFlux: Understanding and Exploiting Layer-Specific Roles in RoPE-Based MMDiT for Versatile Image Editing

ICCV 2025poster

The integration of Rotary Position Embedding (RoPE) in Multimodal Diffusion Transformer (MMDiT) has significantly enhanced text-to-image generation quality. However, the fundamental reliance of self-attention layers on positional embedding versus query-key similarity during generation remains an int…

Cited by 0SourcePDFScholar
2025

GaussianAnything: Interactive Point Cloud Flow Matching for 3D Generation

ICLR 2025poster

Recent advancements in diffusion models and large-scale datasets have revolutionized image and video generation, with increasing focus on 3D content generation. While existing methods show promise, they face challenges in input formats, latent space structures, and output representations. This paper…

Cited by 0SourcePDFScholar
2025

Neural LightRig: Unlocking Accurate Object Normal and Material Estimation with Multi-Light Diffusion

CVPR 2025poster

Recovering the geometry and materials of objects from a single image is challenging due to its under-constrained nature. In this paper, we present Neural LightRig, a novel framework that boosts intrinsic estimation by leveraging auxiliary multi-lighting conditions from 2D diffusion priors. Specifica…

2025

SAR3D: Autoregressive 3D Object Generation and Understanding via Multi-scale 3D VQVAE

CVPR 2025poster

Autoregressive models have demonstrated remarkable success across various fields, from large language models (LLMs) to large multimodal models (LMMs) and 2D content generation, moving closer to artificial general intelligence (AGI). Despite these advances, applying autoregressive approaches to 3D ob…

Cited by 5SourcePDFScholar
2025

Textured 3D Regenerative Morphing with 3D Diffusion Prior

ICCV 2025poster

Textured 3D morphing creates smooth and plausible interpolation sequences between two 3D objects, focusing on transitions in both shape and texture. This is important for creative applications like visual effects in filmmaking. Previous methods rely on establishing point-to-point correspondences and…

2025

TokensGen: Harnessing Condensed Tokens for Long Video Generation

ICCV 2025poster

Generating consistent long videos is a complex challenge: while diffusion-based generative models generate visually impressive short clips, extending them to longer durations often leads to memory bottlenecks and long-term inconsistency. In this paper, we propose TokensGen, a novel two-stage framewo…

Cited by 0SourcePDFScholar
2025

Trajectory attention for fine-grained video motion control

ICLR 2025poster

Recent advancements in video generation have been greatly driven by video diffusion models, with camera motion control emerging as a crucial challenge in creating view-customized visual content. This paper introduces trajectory attention, a novel approach that performs attention along available pixe…

Cited by 0SourcePDFScholar
2025

WorldMem: Long-term Consistent World Simulation with Memory

NeurIPS 2025poster

World simulation has gained increasing popularity due to its ability to model virtual environments and predict the consequences of actions. However, the limited temporal context window often leads to failures in maintaining long-term consistency, particularly in preserving 3D spatial consistency. In…

Cited by 0SourceScholar
2024

ComboVerse: Compositional 3D Assets Creation Using Spatially-Aware Diffusion Guidance

ECCV 2024poster

"Generating high-quality 3D assets from a given image is highly desirable in various applications such as AR/VR. Recent advances in single-image 3D generation explore feed-forward models that learn to infer the 3D model of an object without optimization. Though promising results have been achieved i…

2024

LN3Diff: Scalable Latent Neural Fields Diffusion for Speedy 3D Generation

ECCV 2024poster

"The field of neural rendering has witnessed significant progress with advancements in generative models and differentiable rendering techniques. Though 2D diffusion has achieved success, a unified 3D diffusion pipeline remains unsettled. This paper introduces a novel framework called to address thi…

2024

MVIP-NeRF: Multi-view 3D Inpainting on NeRF Scenes via Diffusion Prior

CVPR 2024poster

Despite the emergence of successful NeRF inpainting methods built upon explicit RGB and depth 2D inpainting supervisions these methods are inherently constrained by the capabilities of their underlying 2D inpainters. This is due to two key reasons: (i) independently inpainting constituent images res…

Cited by 13SourcePDFScholar
2024

Video Diffusion Models are Training-free Motion Interpreter and Controller

NeurIPS 2024poster

Video generation primarily aims to model authentic and customized motion across frames, making understanding and controlling the motion a crucial topic. Most diffusion-based studies on video motion focus on motion customization with training-based paradigms, which, however, demands substantial train…

Cited by 15SourcePDFScholar
2023

AssetField: Assets Mining and Reconfiguration in Ground Feature Plane Representation

ICCV 2023poster

Both indoor and outdoor environments are inherently structured and repetitive. Traditional modeling pipelines keep an asset library storing unique object templates, which is both versatile and memory efficient in practice. Inspired by this observation, we propose AssetField, a novel neural scene rep…

Cited by 12PDFScholar
2023

GlowGAN: Unsupervised Learning of HDR Images from LDR Images in the Wild

ICCV 2023poster

Most in-the-wild images are stored in Low Dynamic Range (LDR) form, serving as a partial observation of the High Dynamic Range (HDR) visual world. Despite limited dynamic range, these LDR images are often captured with different exposures, implicitly containing information about the underlying HDR i…

Cited by 13PDFScholar
2023

Grid-Guided Neural Radiance Fields for Large Urban Scenes

CVPR 2023poster

Purely MLP-based neural radiance fields (NeRF-based methods) often suffer from underfitting with blurred renderings on large-scale scenes due to limited model capacity. Recent approaches propose to geographically divide the scene and adopt multiple sub-NeRFs to model each region individually, leadin…

Cited by 94SourcePDFScholar
2023

Voxurf: Voxel-based Efficient and Accurate Neural Surface Reconstruction

ICLR 2023top-25%

Neural surface reconstruction aims to reconstruct accurate 3D surfaces based on multi-view images. Previous methods based on neural volume rendering mostly train a fully implicit model with MLPs, which typically require hours of training for a single scene. Recent efforts explore the explicit volume…

2022

BungeeNeRF: Progressive Neural Radiance Field for Extreme Multi-Scale Scene Rendering

ECCV 2022poster

"Neural Radiance Field (NeRF) has achieved outstanding performance in modeling 3D objects and controlled scenes, usually under a single scale. In this work, we focus on multi-scale cases where large changes in imagery are observed at drastically different scales. This scenario vastly exists in the r…

Cited by 267SourcePDFScholar
2022

Disentangled3D: Learning a 3D Generative Model With Disentangled Geometry and Appearance From Monocular Images

CVPR 2022poster

Learning 3D generative models from a dataset of monocular images enables self-supervised 3D reasoning and controllable synthesis. State-of-the-art 3D generative models are GANs which use neural 3D volumetric representations for synthesis. Images are synthesized by rendering the volumes from a given…

Cited by 51PDFScholar
2021

A Shading-Guided Generative Implicit Model for Shape-Accurate 3D-Aware Image Synthesis

NeurIPS 2021poster

The advancement of generative radiance fields has pushed the boundary of 3D-aware image synthesis. Motivated by the observation that a 3D object should look realistic from multiple viewpoints, these methods introduce a multi-view constraint as regularization to learn valid 3D radiance fields from 2D…

2021

Do 2D GANs Know 3D Shape? Unsupervised 3D Shape Reconstruction from 2D Image GANs

ICLR 2021oral

Natural images are projections of 3D objects on a 2D image plane. While state-of-the-art 2D generative models like GANs show unprecedented quality in modeling the natural image manifold, it is unclear whether they implicitly capture the underlying 3D object structures. And if so, how could we exploi…

2021

Generative Occupancy Fields for 3D Surface-Aware Image Synthesis

NeurIPS 2021poster

The advent of generative radiance fields has significantly promoted the development of 3D-aware image synthesis. The cumulative rendering process in radiance fields makes training these generative models much easier since gradients are distributed over the entire volume, but leads to diffused object…

2021

Talk-To-Edit: Fine-Grained Facial Editing via Dialog

ICCV 2021poster

Facial editing is an important task in vision and graphics with numerous applications. However, existing works are incapable to deliver a continuous and fine-grained editing mode (e.g., editing a slightly smiling face to a big laughing one) with natural interactions with users. In this work, we prop…

Cited by 141PDFcodeScholar
2020

Channel Equilibrium Networks for Learning Deep Representation

ICML 2020poster

Convolutional Neural Networks (CNNs) are typically constructed by stacking multiple building blocks, each of which contains a normalization layer such as batch normalization (BN) and a rectified linear function such as ReLU. However, this work shows that the combination of normalization and rectifie…

2020

Exploiting Deep Generative Prior for Versatile Image Restoration and Manipulation

ECCV 2020poster

Learning a good image prior is a long-term goal for image restoration and manipulation. While existing methods like deep image prior (DIP) capture low-level image statistics, there are still gaps toward an image prior that captures rich image semantics including color, spatial coherence, textures, a…

2019

Self-Supervised Learning via Conditional Motion Propagation

CVPR 2019poster

Intelligent agent naturally learns from motion. Various self-supervised algorithms have leveraged the motion cues to learn effective visual representations. The hurdle here is that motion is both ambiguous and complex, rendering previous works either suffer from degraded learning efficacy, or resort…

Cited by 63PDFcodeScholar
2019

Switchable Whitening for Deep Representation Learning

ICCV 2019poster

Normalization methods are essential components in convolutional neural networks (CNNs). They either standardize or whiten data using statistics estimated in predefined sets of pixels. Unlike existing works that design normalization techniques for specific tasks, we propose Switchable Whitening (SW),…

Cited by 193PDFcodeScholar
2018

Two at Once: Enhancing Learning and Generalization Capacities via IBN-Net

ECCV 2018poster

Convolutional neural networks (CNNs) have achieved great successes in many computer vision problems. Unlike existing works that designed CNN architectures to improve performance on a single task of a single domain and not generalizable, we present IBN-Net, a novel convolutional architecture, which r…