← Search

Ping Tan

82 accepted papers

2026

MeshFlow: Efficient Artistic Mesh Generation via MeshVAE and Flow-based Diffusion Transformer

CVPR 2026

We present MeshFlow, a new method for compressing and generating artist-like 3D meshes. Current mesh generators often adopt Auto-Regressive (AR) next-token prediction, a natural choice given the discrete nature of mesh connectivity, which, however, scales poorly due to the inference cost being quadr

Cited by 0SourcecodeScholar
2026

PoseMaster: A Unified 3D Native Framework for Stylized Pose Generation

CVPR 2026

Pose stylization, which aims to synthesize stylized content aligning with target poses, serves as a fundamental task across 2D, 3D, and video domains. In the 3D realm, prevailing approaches typically rely on a cascade pipeline: first manipulating the image pose via 2D foundation models and subsequen

Cited by 0SourceScholar
2026

SceneMaker: Open-set 3D Scene Generation with Decoupled De-occlusion and Pose Estimation Model

CVPR 2026

We propose a decoupled 3D scene generation framework called SceneMaker in this work. Due to the lack of sufficient open-set de-occlusion and pose estimation priors, existing methods struggle to simultaneously produce high-quality geometry and accurate poses under severe occlusion and open-set settin

Cited by 0SourcecodeScholar
2026

Switch: Learning Agile Skills Switching for Humanoid Robots

ICRA 2026poster

Recent advancements in whole-body control through deep reinforcement learning have enabled humanoid robots to achieve remarkable progress in real-world challenging locomotion skills. However, existing approaches often struggle with flexible transitions between distinct skills, creating safety concer…

2026

UniTEX: Universal High Fidelity Generative Texturing for 3D Shapes

CVPR 2026

We present UniTEX, a novel two-stage 3D texture generation framework to create high-quality, consistent textures for 3D assets. Existing approaches predominantly rely on UV-based models in the second stage to refine textures after reprojecting the generated multi-view images onto the 3D shapes, whic

Cited by 0SourcecodeScholar
2026

Universal Pose Pretraining for Generalizable Vision-Language-Action Policies

RSS 2026poster

Existing Vision-Language-Action (VLA) models often suffer from feature collapse and low training efficiency because they entangle high-level perception with sparse, embodiment-specific action supervision. Since these models typically rely on VLM backbones optimized for Visual Question Answering (VQA…

Cited by 0SourceScholar
2026

ViLearn: Accelerating Training Convergence of Image-to-3D Generation via Visibility Learning

CVPR 2026

Single-image-to-3D shape generation has seen remarkable progress, driven by latent diffusion models trained on the compressed latent space of 3D VAEs. However, the task remains intrinsically ill-posed: recovering complete 3D geometry--especially occluded surfaces--from a single view is inherently am

Cited by 0SourceScholar
2025

Boost 3D Reconstruction using Diffusion-based Monocular Camera Calibration

ICCV 2025poster

In this paper, we present DM-Calib, a diffusion-based approach for estimating pinhole camera intrinsic parameters from a single input image. Monocular camera calibration is essential for many 3D vision tasks. However, most existing methods depend on handcrafted assumptions or are constrained by limi…

2025

CraftsMan3D: High-fidelity Mesh Generation with 3D Native Diffusion and Interactive Geometry Refiner

CVPR 2025poster

We present a novel generative 3D modeling system, coined CraftsMan, which can generate high-fidelity 3D geometries with highly varied shapes, regular mesh topologies, and detailed surfaces, and, notably, allows for refining the geometry in an interactive manner. Despite the significant advancements…

Cited by 0SourcePDFScholar
2025

Dora: Sampling and Benchmarking for 3D Shape Variational Auto-Encoders

CVPR 2025poster

Recent 3D content generation pipelines commonly employ Variational Autoencoders (VAEs) to encode shapes into compact latent representations for diffusion-based generation. However, the widely adopted uniform point sampling strategy in Shape VAE training often leads to a significant loss of geometric…

2025

MapEval: Towards Unified, Robust and Efficient SLAM Map Evaluation Framework

RA-L 2025

Evaluating massive-scale point cloud maps in Simultaneous Localization and Mapping (SLAM) still remains challenging due to three limitations: lack of unified standards, poor robustness to noise, and computational inefficiency. We propose MapEval, a novel framework for point cloud map assessment. Our

Cited by 17SourcecodeScholar
2025

SkillMimic: Learning Basketball Interaction Skills from Demonstrations

CVPR 2025highlight

Traditional reinforcement learning methods for human-object interaction (HOI) rely on labor-intensive, manually designed skill rewards that do not generalize well across different interactions. We introduce SkillMimic, a unified data-driven framework that fundamentally changes how agents learn inter…

2025

SpatialLM: Training Large Language Models for Structured Indoor Modeling

NeurIPS 2025poster

SpatialLM is a large language model designed to process 3D point cloud data and generate structured 3D scene understanding outputs. These outputs include architectural elements like walls, doors, windows, and oriented object boxes with their semantic categories. Unlike previous methods which exploit…

Cited by 0SourceScholar
2025

SymmCompletion: High-Fidelity and High-Consistency Point Cloud Completion with Symmetry Guidance

AAAI 2025technical

Point cloud completion aims to recover a complete point shape from a partial point cloud. Although existing methods can form satisfactory point clouds in global completeness, they often lose the original geometry details and face the problem of geometric inconsistency between existing point clouds a…

2025

Universal Features Guided Zero-Shot Category-Level Object Pose Estimation

AAAI 2025technical

Object pose estimation, crucial in computer vision and robotics applications, faces challenges with the diversity of unseen categories. We propose a zero-shot method to achieve category-level 6-DOF object pose estimation, which exploits both 2D and 3D universal features of input RGB-D image to estab…

Cited by 0SourcePDFScholar
2024

DVI-SLAM: A Dual Visual Inertial SLAM Network

ICRA 2024poster

Recent deep learning based visual simultaneous localization and mapping (SLAM) methods have made significant progress. However, how to make full use of visual information as well as better integrate with inertial measurement unit (IMU) in visual SLAM has potential research value. This paper proposes…

Cited by 14SourceScholar
2024

Efficient 3D Implicit Head Avatar with Mesh-anchored Hash Table Blendshapes

CVPR 2024poster

3D head avatars built with neural implicit volumetric representations have achieved unprecedented levels of photorealism. However the computational cost of these methods remains a significant barrier to their widespread adoption particularly in real-time applications such as virtual reality and tele…

Cited by 5SourcePDFScholar
2024

Era3D: High-Resolution Multiview Diffusion using Efficient Row-wise Attention

NeurIPS 2024poster

In this paper, we introduce **Era3D**, a novel multiview diffusion method that generates high-resolution multiview images from a single-view image. Despite significant advancements in multiview generation, existing methods still suffer from camera prior mismatch, inefficacy, and low resolution, resu…

Cited by 7SourcePDFScholar
2024

GV-Bench: Benchmarking Local Feature Matching for Geometric Verification of Long-term Loop Closure Detection

IROS 2024poster

Visual loop closure detection is an important module in visual simultaneous localization and mapping (SLAM), which associates current camera observation with previously visited places. Loop closures correct drifts in trajectory estimation to build a globally consistent map. However, a false loop clo…

Cited by 3SourcecodeScholar
2024

GeoWizard: Unleashing the Diffusion Priors for 3D Geometry Estimation from a Single Image

ECCV 2024poster

"∗ Equal contributionWe introduce GeoWizard, a new generative foundation model designed for estimating geometric attributes, , depth and normals, from single images. While significant research has already been conducted in this area, the progress has been substantially limited by the low diversity a…

Cited by 103SourcePDFScholar
2024

PanoContext-Former: Panoramic Total Scene Understanding with a Transformer

CVPR 2024poster

Panoramic images enable deeper understanding and more holistic perception of 360 surrounding environment which can naturally encode enriched scene context information compared to standard perspective image. Previous work has made lots of effort to solve the scene understanding task in a hybrid solut…

Cited by 11SourcePDFScholar
2024

PointRegGPT: Boosting 3D Point Cloud Registration using Generative Point-Cloud Pairs for Training

ECCV 2024poster

"Data plays a crucial role in training learning-based methods for 3D point cloud registration. However, the real-world dataset is expensive to build, while rendering-based synthetic data suffers from domain gaps. In this work, we present , boosting 3D Point cloud Registration using Generative Point-…

2024

SweetDreamer: Aligning Geometric Priors in 2D diffusion for Consistent Text-to-3D

ICLR 2024poster

It is inherently ambiguous to lift 2D results from pre-trained diffusion models to a 3D world for text-to-3D generation. 2D diffusion models solely learn view-agnostic priors and thus lack 3D knowledge during the lifting, leading to the multi-view inconsistency problem. We find that this problem pri…

2024

VMINer: Versatile Multi-view Inverse Rendering with Near- and Far-field Light Sources

CVPR 2024highlight

This paper introduces a versatile multi-view inverse rendering framework with near- and far-field light sources. Tackling the fundamental challenge of inherent ambiguity in inverse rendering our framework adopts a lightweight yet inclusive lighting model for different near- and far-field lights thus…

Cited by 0SourcePDFScholar
2023

DENSE RGB SLAM WITH NEURAL IMPLICIT MAPS

ICLR 2023poster

There is an emerging trend of using neural implicit functions for map representation in Simultaneous Localization and Mapping (SLAM). Some pioneer works have achieved encouraging results on RGB-D SLAM. In this paper, we present a dense RGB SLAM method with neural implicit map representation. To reac…

2023

DPS-Net: Deep Polarimetric Stereo Depth Estimation

ICCV 2023poster

Stereo depth estimation usually struggles to deal with textureless scenes for both traditional and learning-based methods due to the inherent dependence on image correspondence matching. In this paper, we propose a novel neural network, i.e., DPS-Net, to exploit both the prior geometric knowledge an…

Cited by 28PDFScholar
2023

DRO: Deep Recurrent Optimizer for Video to Depth

RA-L 2023

There are increasing interests of studying the video-to-depth (V2D) problem with machine learning techniques. While earlier methods directly learn a mapping from images to depth maps and camera poses, more recent works enforce multi-view geometry constraints through optimization embedded in the lear

Cited by 21SourcecodeScholar
2023

Learning Optical Flow from Event Camera with Rendered Dataset

ICCV 2023poster

We study the problem of estimating optical flow from event cameras. One important issue is how to build a high-quality event-flow dataset with accurate event values and flow labels. Previous datasets are created by either capturing real scenes by event cameras or synthesizing from images with pasted…

Cited by 19PDFcodeScholar
2023

Learning Personalized High Quality Volumetric Head Avatars From Monocular RGB Videos

CVPR 2023poster

We propose a method to learn a high-quality implicit 3D head avatar from a monocular RGB video captured in the wild. The learnt avatar is driven by a parametric face model to achieve user-controlled facial expressions and head poses. Our hybrid pipeline combines the geometry prior and dynamic tracki…

Cited by 20SourcePDFScholar
2023

NeuMap: Neural Coordinate Mapping by Auto-Transdecoder for Camera Localization

CVPR 2023poster

This paper presents an end-to-end neural mapping method for camera localization, dubbed NeuMap, encoding a whole scene into a grid of latent codes, with which a Transformer-based auto-decoder regresses 3D coordinates of query pixels. State-of-the-art feature matching methods require each scene to be…

2022

A Real World Dataset for Multi-View 3D Reconstruction

ECCV 2022poster

"We present a dataset of 371 3D models of everyday tabletop objects along with their 320,000 real world RGB and depth images. Accurate annotations of camera poses and object poses for each image are performed in a semi-automated fashion to facilitate the use of the dataset for myriad 3D applications…

2022

A Speech-driven Sign Language Avatar Animation System for Hearing Impaired Applications

IJCAI 2022poster

Sign language is the communication language used in hearing impaired community. Recently, the research of sign language production has made great progress but still need to cope with some critical challenges. In this paper, we propose a system-level scheme and push forward the implementation of sign…

Cited by 6SourcePDFScholar
2022

AutoAvatar: Autoregressive Neural Fields for Dynamic Avatar Modeling

ECCV 2022poster

"Neural fields such as implicit surfaces have recently enabled avatar modeling from raw scans without explicit temporal correspondences. In this work, we exploit autoregressive modeling to further extend this notion to capture dynamic effects, such as soft-tissue deformations. Although autoregressiv…

2022

DART: Articulated Hand Model with Diverse Accessories and Rich Textures

NeurIPS 2022accept

Hand, the bearer of human productivity and intelligence, is receiving much attention due to the recent fever of digital twins. Among different hand morphable models, MANO has been widely used in vision and graphics community. However, MANO disregards textures and accessories, which largely limits it…

2022

Domain Randomization-Enhanced Depth Simulation and Restoration for Perceiving and Grasping Specular and Transparent Objects

ECCV 2022poster

"Commercial depth sensors usually generate noisy and missing depths, especially on specular and transparent objects, which poses critical issues to downstream depth or point cloud-based tasks. To mitigate this problem, we propose a powerful RGBD fusion network, SwinDRNet, for depth restoration. We f…

2022

Efficient Virtual View Selection for 3D Hand Pose Estimation

AAAI 2022technical

3D hand pose estimation from single depth is a fundamental problem in computer vision, and has wide applications. However, the existing methods still can not achieve satisfactory hand pose estimation results due to view variation and occlusion of human hand. In this paper, we propose a new virtual v…

2022

GAT-CADNet: Graph Attention Network for Panoptic Symbol Spotting in CAD Drawings

CVPR 2022poster

Spotting graphical symbols from the computer-aided design (CAD) drawings is essential to many industrial applications. Different from raster images, CAD drawings are vector graphics consisting of geometric primitives such as segments, arcs, and circles. By treating each CAD drawing as a graph, we pr…

Cited by 18PDFScholar
2022

Neural Window Fully-Connected CRFs for Monocular Depth Estimation

CVPR 2022poster

Estimating the accurate depth from a single image is challenging since it is inherently ambiguous and ill-posed. While recent works design increasingly complicated and powerful networks to directly regress the depth map, we take the path of CRFs optimization. Due to the expensive computation, CRFs a…

Cited by 424PDFScholar
2022

OCRTOC: A Cloud-Based Competition and Benchmark for Robotic Grasping and Manipulation

RA-L 2022

In this paper, we propose a cloud-based benchmark for robotic grasping and manipulation, called the OCRTOC benchmark. The benchmark focuses on the object rearrangement problem, specifically table organization tasks. We provide a set of identical real robot setups and facilitate remote experiments of

Cited by 58SourcecodeScholar
2022

RCP: Recurrent Closest Point for Point Cloud

CVPR 2022oral

3D motion estimation including scene flow and point cloud registration has drawn increasing interest. Inspired by 2D flow estimation, recent methods employ deep neural networks to construct the cost volume for estimating accurate 3D flow. However, these methods are limited by the fact that it is dif…

Cited by 34PDFcodeScholar
2022

SceneSqueezer: Learning To Compress Scene for Camera Relocalization

CVPR 2022oral

Standard visual localization methods build a priori 3D model of a scene which is used to establish correspondences against the 2D keypoints in a query image. Storing these pre-built 3D scene models can be prohibitively expensive for large-scale environments, especially on mobile devices with limited…

Cited by 37PDFScholar
2022

Streaming Radiance Fields for 3D Video Synthesis

NeurIPS 2022accept

We present an explicit-grid based method for efficiently reconstructing streaming radiance fields for novel view synthesis of real world dynamic scenes. Instead of training a single model that combines all the frames, we formulate the dynamic modeling problem with an incremental learning paradigm in…

2021

CondLaneNet: A Top-To-Down Lane Detection Framework Based on Conditional Convolution

ICCV 2021poster

Modern deep-learning-based lane detection methods are successful in most scenarios but struggling for lane lines with complex topologies. In this work, we propose CondLaneNet, a novel top-to-down lane detection framework that detects the lane instances first and then dynamically predicts the line sh…

Cited by 315PDFcodeScholar
2021

End-to-End Rotation Averaging With Multi-Source Propagation

CVPR 2021poster

This paper presents an end-to-end neural network for multiple rotation averaging in SfM. Due to the manifold constraint of rotations, conventional methods usually take two separate steps involving spanning tree based initialization and iterative nonlinear optimization respectively. These methods can…

Cited by 29PDFcodeScholar
2021

FloorPlanCAD: A Large-Scale CAD Drawing Dataset for Panoptic Symbol Spotting

ICCV 2021poster

Access to large and diverse computer-aided design (CAD) drawings is critical for developing symbol spotting algorithms. In this paper, we present FloorPlanCAD, a large-scale real-world CAD drawing dataset containing over 10,000 floor plans, ranging from residential to commercial buildings. CAD drawi…

Cited by 50PDFcodeScholar
2021

HumanGPS: Geodesic PreServing Feature for Dense Human Correspondences

CVPR 2021poster

In this paper, we address the problem of building pixel-wise dense correspondences between human images under arbitrary camera viewpoints and body poses. Previous methods either assume small motions or rely on discriminative descriptors extracted from local patches, which cannot handle large motion…

Cited by 14PDFScholar
2021

Interacting Two-Hand 3D Pose and Shape Reconstruction From Single Color Image

ICCV 2021poster

In this paper, we propose a novel deep learning framework to reconstruct 3D hand poses and shapes of two interacting hands from a single color image. Previous methods designed for single hand cannot be easily applied for the two hand scenario because of the heavy inter-hand occlusion and larger solu…

Cited by 111PDFcodeScholar
2021

Learning Efficient Photometric Feature Transform for Multi-View Stereo

ICCV 2021poster

We present a novel framework to learn to convert the per-pixel photometric information at each view into spatially distinctive and view-invariant low-level features, which can be plugged into existing multi-view stereo pipeline for enhanced 3D reconstruction. Both the illumination conditions during…

Cited by 3PDFScholar
2021

Single-Shot is Enough: Panoramic Infrastructure Based Calibration of Multiple Cameras and 3D LiDARs

IROS 2021poster

The integration of multiple cameras and 3D Li-DARs has become basic configuration of augmented reality devices, robotics, and autonomous vehicles. The calibration of multi-modal sensors is crucial for a system to properly function, but it remains tedious and impractical for mass production. Moreover…

Cited by 27SourceScholar
2021

Stereo Matching by Self-supervision of Multiscopic Vision

IROS 2021poster

Self-supervised learning for depth estimation possesses several advantages over supervised learning. The benefits of no need for ground-truth depth, online fine-tuning, and better generalization with unlimited data attract researchers to seek self-supervised solutions. In this work, we propose a new…

Cited by 18SourceScholar
2020

Cascade Cost Volume for High-Resolution Multi-View Stereo and Stereo Matching

CVPR 2020oral

The deep multi-view stereo (MVS) and stereo matching approaches generally construct 3D cost volumes to regularize and regress the output depth or disparity. These methods are limited when high-resolution outputs are needed since the memory and time costs grow cubically as the volume resolution incre…

Cited by 897PDFcodeScholar
2020

Channel Equilibrium Networks for Learning Deep Representation

ICML 2020poster

Convolutional Neural Networks (CNNs) are typically constructed by stacking multiple building blocks, each of which contains a normalization layer such as batch normalization (BN) and a rectified linear function such as ReLU. However, this work shows that the combination of normalization and rectifie…

2020

End-to-End Learning Local Multi-View Descriptors for 3D Point Clouds

CVPR 2020poster

In this work, we propose an end-to-end framework to learn local multi-view descriptors for 3D point clouds. To adopt a similar multi-view representation, existing studies use hand-crafted viewpoints for rendering in a preprocessing stage, which is detached from the subsequent descriptor learning sta…

Cited by 144PDFScholar
2020

Interpretable Foreground Object Search As Knowledge Distillation

ECCV 2020poster

This paper proposes a knowledge distillation method for foreground object search (FoS). Given a background and a rectangle specifying the foreground location and scale, FoS retrieves compatible foregrounds in a certain category for later image composition. Foregrounds within the same category can be…

Cited by 7SourcePDFScholar
2020

Self-Supervised Human Depth Estimation From Monocular Videos

CVPR 2020poster

Previous methods on estimating detailed human depth often require supervised training with 'ground truth' depth data. This paper presents a self-supervised method that can be trained on YouTube videos without known depth, which makes training data collection simple and improves the generalization of…

Cited by 35PDFScholar
2019

A Neural Network for Detailed Human Depth Estimation From a Single Image

ICCV 2019oral

This paper presents a neural network to estimate a detailed depth map of the foreground human in a single RGB image. The result captures geometry details such as cloth wrinkles, which are important in visualization applications. To achieve this goal, we separate the depth map into a smooth base shap…

Cited by 60PDFcodeScholar
2019

Batch DropBlock Network for Person Re-Identification and Beyond

ICCV 2019poster

Since the person re-identification task often suffers from the problem of pose changes and occlusions, some attentive local features are often suppressed when training CNNs. In this paper, we propose the Batch DropBlock (BDB) Network which is a two branch network composed of a conventional ResNet-50…

Cited by 317PDFScholar
2019

Learned Map Prediction for Enhanced Mobile Robot Exploration

ICRA 2019poster

We demonstrate an autonomous ground robot capable of exploring unknown indoor environments for reconstructing their 2D maps. This problem has been traditionally tackled by geometric heuristics and information theory. More recently, deep learning and reinforcement learning based approaches have been…

Cited by 123SourceScholar
2019

SANet: Scene Agnostic Network for Camera Localization

ICCV 2019poster

This paper presents a scene agnostic neural architecture for camera localization, where model parameters and scenes are independent from each other.Despite recent advancement in learning based methods, most approaches require training for each scene one by one, not applicable for online applications…

Cited by 99PDFScholar
2018

Faces as Lighting Probes via Unsupervised Deep Highlight Extraction

ECCV 2018poster

We present a method for estimating detailed scene illumination using human faces in a single image. In contrast to previous works that estimate lighting in terms of low-order basis functions or distant point lights, our technique estimates illumination at a higher precision in the form of a non-para…

Cited by 52SourcePDFScholar
2018

Sparsely Aggregated Convolutional Networks

ECCV 2018poster

We explore a key architectural aspect of deep convolutional neural networks: the pattern of internal skip connections used to aggregate outputs of earlier layers for consumption by deeper layers. Such aggregation is critical to facilitate training of very deep networks in an end-to-end manner. This…

2018

Very Large-Scale Global SfM by Distributed Motion Averaging

CVPR 2018poster

Global Structure-from-Motion (SfM) techniques have demonstrated superior efficiency and accuracy than the conventional incremental approach in many recent studies. This work proposes a divide-and-conquer framework to solve very large global SfM at the scale of millions of images. Specifically, we fi…

Cited by 183SourcePDFScholar
2016

A Benchmark Dataset and Evaluation for Non-Lambertian and Uncalibrated Photometric Stereo

CVPR 2016poster

Recent progress on photometric stereo extends the technique to deal with general materials and unknown illumination conditions. However, due to the lack of suitable benchmark data with ground truth shapes (normals), quantitative comparison and evaluation is difficult to achieve. In this paper, we fi…

Cited by 358PDFScholar
2015

Simultaneous Video Defogging and Stereo Reconstruction

CVPR 2015poster

We present a method to jointly estimate scene depth and recover the clear latent image from a foggy video sequence. In our formulation, the depth cues from stereo matching and fog information reinforce each other, and produce superior results than conventional stereo or defogging algorithms. We firs…

Cited by 166SourcePDFScholar