← Search

Fangneng Zhan

34 accepted papers

2026

AREA3D: Active Reconstruction Agent with Unified Feed-Forward 3D Perception and Vision-Language Guidance

CVPR 2026

Active 3D reconstruction enables an agent to autonomously select viewpoints to build accurate and complete scene geometry efficiently, rather than passively reconstructing scenes from pre-collected images. Existing active reconstruction methods often rely on geometric heuristics, which may result in

Cited by 0SourcecodeScholar
2026

Abstract 3D Perception for Spatial Intelligence in Vision-Language Models

CVPR 2026

Vision-language models (VLMs) struggle with 3D-related tasks such as spatial cognition and physical understanding, which are crucial for real-world applications like robotics and embodied agents. We attribute this to a modality gap between the 3D tasks and the 2D training of VLM, which led to ineffi

Cited by 0SourceScholar
2026

Flow Equivariant World Models: Structured Memory for Dynamic Environments

ICML 2026poster

The natural world is richly structured over space and time. Much of this structure arises from the interplay between spatial geometry and motion. However, most existing world models ignore this structure, leading to an inability to generalize in dynamic environments. In this work, we show that enfor…

Cited by 0SourcecodeScholar
2026

MuSASplat: Efficient Sparse-View 3D Gaussian Splats via Lightweight Multi-Scale Adaptation

AAAI 2026technical

Sparse-view 3D Gaussian splatting seeks to render high-quality novel views of 3D scenes from a limited set of input images. While recent pose-free feed-forward methods leveraging pre-trained 3D priors have achieved impressive results, most of them rely on full fine-tuning of large Vision Transformer

Cited by 0SourcePDFScholar
2026

PAGE-4D: Disentangled Pose and Geometry Estimation for 4D Perception

ICLR 2026poster

Recent 3D feed-forward models, such as the Visual Geometry Grounded Transformer (VGGT), have shown strong capability in inferring 3D attributes of static scenes. However, since they are typically trained on static datasets, these models often struggle in real-world scenarios involving complex dynami…

Cited by 27SourcecodeScholar
2026

RoboTAG: End-to-end Robot Pose Estimation via Topological Alignment Graph

CVPR 2026

Estimating robot pose from a monocular RGB image is a challenge in robotics and computer vision. Existing methods typically build networks on top of 2D visual backbones and depend heavily on labeled data for training, which is often scarce in real-world scenarios, causing a sim-to-real gap. Moreover

Cited by 0SourceScholar
2025

VidSeg: Training-free Video Semantic Segmentation based on Diffusion Models

CVPR 2025poster

We introduce the first training-free approach for Video Semantic Segmentation (VSS) based on pre-trained diffusion models. A growing research direction attempts to employ diffusion models to perform downstream vision tasks by exploiting their deep understanding of image semantics. Yet, the majority…

Cited by 0SourcePDFScholar
2024

DatasetNeRF: Efficient 3D-aware Data Factory with Generative Radiance Fields

ECCV 2024poster

"Progress in 3D computer vision tasks demands a huge amount of data, yet annotating multi-view images with 3D-consistent annotations, or point clouds with part segmentation is both time-consuming and challenging. This paper introduces DatasetNeRF, a novel approach capable of generating infinite, hig…

2024

FreGS: 3D Gaussian Splatting with Progressive Frequency Regularization

CVPR 2024poster

3D Gaussian splatting has achieved very impressive performance in real-time novel view synthesis. However it often suffers from over-reconstruction during Gaussian densification where high-variance image regions are covered by a few large Gaussians only leading to blur and artifacts in the rendered…

Cited by 55SourcePDFScholar
2023

KD-DLGAN: Data Limited Image Generation via Knowledge Distillation

CVPR 2023poster

Generative Adversarial Networks (GANs) rely heavily on large-scale training data for training high-quality image generation models. With limited training data, the GAN discriminator often suffers from severe overfitting which directly leads to degraded generation especially in generation diversity.…

Cited by 29SourcePDFScholar
2023

Pose-Free Neural Radiance Fields via Implicit Pose Regularization

ICCV 2023poster

Pose-free neural radiance fields (NeRF) aim to train NeRF with unposed multi-view images and it has achieved very impressive success in recent years. Most existing works share the pipeline of training a coarse pose estimator with rendered images at first, followed by a joint optimization of estimate…

Cited by 11PDFScholar
2023

Regularized Vector Quantization for Tokenized Image Synthesis

CVPR 2023poster

Quantizing images into discrete representations has been a fundamental problem in unified generative modeling. Predominant approaches learn the discrete representation either in a deterministic manner by selecting the best-matching token or in a stochastic manner by sampling from a predicted distrib…

2023

StyleRF: Zero-Shot 3D Style Transfer of Neural Radiance Fields

CVPR 2023poster

3D style transfer aims to render stylized novel views of a 3D scene with multi-view consistency. However, most existing work suffers from a three-way dilemma over accurate geometry reconstruction, high-quality stylization, and being generalizable to arbitrary new styles. We propose StyleRF (Style Ra…

2023

WaveNeRF: Wavelet-based Generalizable Neural Radiance Fields

ICCV 2023poster

Neural Radiance Field (NeRF) has shown impressive performance in novel view synthesis via implicit scene representation. However, it usually suffers from poor scalability as requiring densely sampled images for each new scene. Several studies have attempted to mitigate this problem by integrating Mu…

Cited by 16PDFScholar
2023

Weakly Supervised 3D Open-vocabulary Segmentation

NeurIPS 2023poster

Open-vocabulary segmentation of 3D scenes is a fundamental function of human perception and thus a crucial objective in computer vision research. However, this task is heavily impeded by the lack of large-scale and diverse 3D open-vocabulary segmentation datasets for training robust and generalizabl…

2022

Auto-Regressive Image Synthesis with Integrated Quantization

ECCV 2022poster

"Deep generative models have achieved conspicuous progress in realistic image synthesis with multifarious conditional inputs, while generating diverse yet high-fidelity images remains a grand challenge in conditional image generation. This paper presents a versatile framework for conditional image g…

2022

Bi-Level Feature Alignment for Versatile Image Translation and Manipulation

ECCV 2022poster

"Generative adversarial networks (GANs) have achieved great success in image translation and manipulation. However, high-fidelity image generation with faithful style control remains a grand challenge in computer vision. This paper presents a versatile image translation and manipulation framework th…

Cited by 52SourcePDFScholar
2022

Fourier Document Restoration for Robust Document Dewarping and Recognition

CVPR 2022poster

State-of-the-art document dewarping techniques learn to predict 3-dimensional information of documents which are prone to errors while dealing with documents with irregular distortions or large variations in depth. This paper presents FDRNet, a Fourier Document Restoration Network that can restore d…

Cited by 32PDFcodeScholar
2022

GenCo: Generative Co-training for Generative Adversarial Networks with Limited Data

AAAI 2022technical

Training effective Generative Adversarial Networks (GANs) requires large amounts of training data, without which the trained models are usually sub-optimal with discriminator over-fitting. Several prior studies address this issue by expanding the distribution of the limited training data via massive…

Cited by 39SourcePDFScholar
2022

Marginal Contrastive Correspondence for Guided Image Generation

CVPR 2022oral

Exemplar-based image translation establishes dense correspondences between a conditional input and an exemplar (from two different domains) for leveraging detailed exemplar styles to achieve realistic image translation. Existing work builds the cross-domain correspondences implicitly by minimizing f…

Cited by 78PDFScholar
2022

Masked Generative Adversarial Networks are Data-Efficient Generation Learners

NeurIPS 2022accept

This paper shows that masked generative adversarial network (MaskedGAN) is robust image generation learners with limited training data. The idea of MaskedGAN is simple: it randomly masks out certain image information for effective GAN training with limited data. We develop two masking strategies tha…

Cited by 28SourcePDFScholar
2022

Modulated Contrast for Versatile Image Synthesis

CVPR 2022poster

Perceiving the similarity between images has been a long-standing and fundamental problem underlying various visual generation tasks. Predominant approaches measure the inter-image distance by computing pointwise absolute deviations, which tends to estimate the median of instance distributions and l…

Cited by 216PDFcodeScholar
2022

Transfer Learning from Synthetic to Real LiDAR Point Cloud for Semantic Segmentation

AAAI 2022technical

Knowledge transfer from synthetic to real data has been widely studied to mitigate data annotation constraints in various computer vision tasks such as semantic segmentation. However, the study focused on 2D images and its counterpart in 3D point clouds segmentation lags far behind due to the lack o…

2021

EMLight: Lighting Estimation via Spherical Distribution Approximation

AAAI 2021technical

Illumination estimation from a single image is critical in 3D rendering and it has been investigated extensively in the computer vision and computer graphic research community. On the other hand, existing works estimate illumination by either regressing light parameters or generating illumination ma…

Cited by 165SourcePDFScholar
2021

Sparse Needlets for Lighting Estimation With Spherical Transport Loss

ICCV 2021poster

Accurate lighting estimation is challenging yet critical to many computer vision and computer graphics tasks such as high-dynamic-range (HDR) relighting. Existing approaches model lighting in either frequency domain or spatial domain which is insufficient to represent the complex lighting conditions…

Cited by 112PDFScholar
2021

Unbalanced Feature Transport for Exemplar-Based Image Translation

CVPR 2021poster

Despite the great success of GANs in images translation with different conditioned inputs such as semantic segmentation and edge map, generating high-fidelity images with reference styles from exemplars remains a grand challenge in conditional image-to-image translation. This paper presents a genera…

Cited by 235PDFScholar
2021

WaveFill: A Wavelet-Based Generation Network for Image Inpainting

ICCV 2021poster

Image inpainting aims to complete the missing or corrupted regions of images with realistic contents. The prevalent approaches adopt a hybrid objective of reconstruction and perceptual quality by using generative adversarial networks. However, the reconstruction loss and adversarial loss focus on sy…

Cited by 132PDFcodeScholar
2019

GA-DAN: Geometry-Aware Domain Adaptation Network for Scene Text Detection and Recognition

ICCV 2019poster

Recent adversarial learning research has achieved very impressive progress for modelling cross-domain data shifts in appearance space but its counterpart in modelling cross-domain shifts in geometry space lags far behind. This paper presents an innovative Geometry-Aware Domain Adaptation Network (GA…

Cited by 350PDFScholar
2018

Accurate Scene Text Detection through Border Semantics Awareness and Bootstrapping

ECCV 2018poster

This paper presents a scene text detection technique that exploits bootstrapping and text border semantics for accurate localization of texts in scenes. A novel bootstrapping technique is designed which samples multiple ‘subsections’ of a word or text line and accordingly relieves the constraint of…

Cited by 148SourcePDFScholar
2018

Verisimilar Image Synthesis for Accurate Detection and Recognition of Texts in Scenes

ECCV 2018poster

The requirement of large amounts of annotated images has become one grand challenge while training deep neural network models for various visual detection and recognition tasks. This paper presents a novel image synthesis technique that aims to generate a large amount of annotated scene text images…

Cited by 341SourcePDFScholar