← Search

Qixing Huang

78 accepted papers

2026

Decoupled Generative Modeling for Human-Object Interaction Synthesis

CVPR 2026

Synthesizing realistic human-object interaction (HOI) is essential for 3D computer vision and robotics, underpinning animation and embodied control. Existing approaches often require manually specified intermediate waypoints and place all optimization objectives on a single network, which increases

Cited by 0SourceScholar
2026

LiteGE: Lightweight Geodesic Embedding for Efficient Geodesics Computation and Non-Isometric Shape Correspondence

AAAI 2026technical

Computing geodesic distances on 3D surfaces is fundamental to many tasks in 3D vision and geometry processing, with deep connections to tasks such as shape correspondence. Recent learning-based methods achieve strong performance but rely on large 3D backbones, leading to high memory usage and latenc

Cited by 0SourcePDFScholar
2026

Mining Attribute Subspaces for Efficient Fine-tuning of 3D Foundation Models

CVPR 2026

With the emergence of 3D foundation models, there is growing interest in fine-tuning them for downstream tasks, where LoRA is the dominant fine-tuning paradigm. As 3D datasets exhibit distinct variations in texture, geometry, camera motion, and lighting, there are interesting fundamental questions:

Cited by 0SourceScholar
2026

Recovering Physically Plausible Human-Object Interactions from Monocular Videos

CVPR 2026

In this paper, we present a method to reconstruct physically plausible human-object interactions (HOI) from monocular videos. While existing kinematic-based approaches produce visually plausible motion, they often result in physical artifacts such as interpenetration and object floating. To overcome

Cited by 0SourceScholar
2026

Revisiting Spectral Representations in Generative Diffusion Models

ICML 2026poster

Diffusion models have shown remarkable performance on diverse generation tasks. Recent work finds that imposing representation alignment on the hidden states of diffusion networks can both facilitate training convergence and enhance sampling quality, yet the mechanism driving this synergy remains in…

Cited by 0SourceScholar
2026

SounDiT: Geo-Contextual Soundscape-to-Landscape Generation

CVPR 2026

Recent audio-to-image models have shown impressive performance in generating images of specific objects conditioned on their corresponding sounds. However, these models fail to reconstruct real-world landscapes conditioned on acoustic environments. To address this challenge, we present Geo-contextua

Cited by 0SourceScholar
2026

WorldReel: 4D Video Generation with Consistent Geometry and Motion Modeling

CVPR 2026

Recent video generators achieve striking photorealism, yet remain fundamentally inconsistent in 3D. We present WorldReel, a 4D video generator that is natively spatio-temporally consistent. WorldReel jointly produces RGB frames together with 4D scene representations, including pointmaps, camera traj

Cited by 0SourcecodeScholar
2025

Atlas Gaussians Diffusion for 3D Generation

ICLR 2025spotlight

Using the latent diffusion model has proven effective in developing novel 3D generation techniques. To harness the latent diffusion model, a key challenge is designing a high-fidelity and efficient representation that links the latent space and the 3D space. In this paper, we introduce Atlas Gaussia…

2025

CADCrafter: Generating Computer-Aided Design Models from Unconstrained Images

CVPR 2025poster

Creating CAD digital twins from the physical world is crucial for manufacturing, design, and simulation. However, current methods typically rely on costly 3D scanning with labor-intensive post-processing. To provide a user-friendly design process, we explore the problem of reverse engineering from u…

Cited by 3SourcePDFScholar
2025

Escaping Plato's Cave: Towards the Alignment of 3D and Text Latent Spaces

CVPR 2025poster

Recent works have shown that, when trained at scale, uni-modal 2D vision and text encoders converge to learned features that share remarkable structural properties, despite arising from different representations. However, the role of 3D encoders with respect to other modalities remains unexplored. F…

Cited by 0SourcePDFScholar
2025

FiffDepth: Feed-forward Transformation of Diffusion-Based Generators for Detailed Depth Estimation

ICCV 2025poster

Monocular Depth Estimation (MDE) is a fundamental 3D vision problem with numerous applications such as 3D scene reconstruction, autonomous navigation, and AI content creation. However, robust and generalizable MDE remains challenging due to limited real-world labeled data and distribution gaps betwe…

Cited by 0SourcePDFScholar
2025

GenVDM: Generating Vector Displacement Maps From a Single Image

CVPR 2025highlight

We introduce the first method for generating Vector Displacement Maps (VDMs): parameterized, detailed geometric stamps commonly used in 3D modeling. Given a single input image, our method first generates multi-view normal maps and then reconstructs a VDM from the normals via a novel reconstruction p…

2025

GeoVideo: Introducing Geometric Regularization into Video Generation Model

NeurIPS 2025poster

Recent advances in video generation have enabled the synthesis of high-quality and visually realistic clips using diffusion transformer models. However, most existing approaches operate purely in the 2D pixel space and lack explicit mechanisms for modeling 3D structures, often resulting in temporall…

Cited by 0SourceScholar
2025

HUMOTO: A 4D Dataset of Mocap Human Object Interactions

ICCV 2025poster

We present Human Motions with Objects (HUMOTO), a high-fidelity dataset of human-object interactions for motion generation, computer vision, and robotics applications. Featuring 735 sequences (7,875 seconds at 30 fps), HUMOTO captures interactions with 63 precisely modeled objects and 72 articulated…

2025

MegaSynth: Scaling Up 3D Scene Reconstruction with Synthesized Data

CVPR 2025poster

We propose scaling up 3D scene reconstruction by training with synthesized data. At the core of our work is MegaSynth, a procedurally generated 3D dataset comprising 700K scenes - over 50 times larger than the prior real dataset DL3DV - dramatically scaling the training data. To enable scalable data…

Cited by 1SourcePDFScholar
2025

PASTA: Part-Aware Sketch-to-3D Shape Generation with Text-Aligned Prior

ICCV 2025poster

A fundamental challenge in conditional 3D shape generation is to minimize the information loss and maximize the intention of user input. Existing approaches have predominantly focused on two types of isolated conditional signals, i.e., user sketches and text descriptions, each of which does not offe…

Cited by 0SourcePDFScholar
2025

RayZer: A Self-supervised Large View Synthesis Model

ICCV 2025poster

We present RayZer, a self-supervised multi-view 3D Vision model trained without any 3D supervision, i.e., camera poses and scene geometry, while exhibiting emerging 3D awareness. Concretely, RayZer takes unposed and uncalibrated images as input, recovers camera parameters, reconstructs a scene repre…

Cited by 0SourcePDFScholar
2025

Real3D: Towards Scaling Large Reconstruction Models with Real Images

ICCV 2025poster

Training single-view Large Reconstruction Models (LRMs) follows the fully supervised route, requiring multi-view supervision. However, the multi-view data typically comes from synthetic 3D assets, which are hard to scale further and are not representative of the distribution of real-world object sha…

Cited by 0SourcePDFScholar
2025

Reconstructing Humans with a Biomechanically Accurate Skeleton

CVPR 2025poster

In this paper, we introduce a method for reconstructing 3D humans from a single image using a biomechanically accurate skeleton model. To achieve this, we train a transformer that takes an image as input and estimates the parameters of the model. Due to the lack of training data for this task, we bu…

2024

3D Feature Prediction for Masked-AutoEncoder-Based Point Cloud Pretraining

ICLR 2024poster

Masked autoencoders (MAE) have recently been introduced to 3D self-supervised pretraining for point clouds due to their great success in NLP and computer vision. Unlike MAEs used in the image domain, where the pretext task is to restore features at the masked pixels, such as colors, the existing 3D…

2024

CoFie: Learning Compact Neural Surface Representations with Coordinate Fields

NeurIPS 2024poster

This paper introduces CoFie, a novel local geometry-aware neural surface representation. CoFie is motivated by the theoretical analysis of local SDFs with quadratic approximation. We find that local shapes are highly compressive in an aligned coordinate frame defined by the normal and tangent direct…

2024

DeblurSR: Event-Based Motion Deblurring under the Spiking Representation

AAAI 2024technical

We present DeblurSR, a novel motion deblurring approach that converts a blurry image into a sharp video. DeblurSR utilizes event data to compensate for motion ambiguities and exploits the spiking representation to parameterize the sharp output video as a mapping from time to intensity. Our key contr…

2024

Detector-Free Structure from Motion

CVPR 2024poster

We propose a structure-from-motion framework to recover accurate camera poses and point clouds from unordered images. Traditional SfM systems typically rely on the successful detection of repeatable keypoints across multiple views as the first step which is difficult for texture-poor scenes and poor…

2024

Enhancing Implicit Shape Generators Using Topological Regularizations

ICML 2024poster

A fundamental problem in learning 3D shapes generative models is that when the generative model is simply fitted to the training data, the resulting synthetic 3D models can present various artifacts. Many of these artifacts are topological in nature, e.g., broken legs, unrealistic thin structures, a…

Cited by 1SourcePDFScholar
2024

GPLD3D: Latent Diffusion of 3D Shape Generative Models by Enforcing Geometric and Physical Priors

CVPR 2024poster

State-of-the-art man-made shape generative models usually adopt established generative models under a suitable implicit shape representation. A common theme is to perform distribution alignment which does not explicitly model important shape priors. As a result many synthetic shapes are not connecte…

Cited by 8SourcePDFScholar
2024

GenCorres: Consistent Shape Matching via Coupled Implicit-Explicit Shape Generative Models

ICLR 2024poster

This paper introduces GenCorres, a novel unsupervised joint shape matching (JSM) approach. Our key idea is to learn a mesh generator to fit an unorganized deformable shape collection while constraining deformations between adjacent synthetic shapes to preserve geometric structures such as local rigi…

2024

High-Fidelity and Transferable NeRF Editing by Frequency Decomposition

ECCV 2024poster

"This paper enables high-fidelity, transferable NeRF editing by frequency decomposition. Recent NeRF editing pipelines lift 2D stylization results to 3D scenes while suffering from blurry results, and fail to capture detailed structures caused by the inconsistency between 2D editings. Our critical i…

2024

LEAP: Liberate Sparse-View 3D Modeling from Camera Poses

ICLR 2024poster

Are camera poses necessary for multi-view 3D modeling? Existing approaches predominantly assume access to accurate camera poses. While this assumption might hold for dense views, accurately estimating camera poses for sparse views is often elusive. Our analysis reveals that noisy estimated poses lea…

2024

MoGenTS: Motion Generation based on Spatial-Temporal Joint Modeling

NeurIPS 2024poster

Motion generation from discrete quantization offers many advantages over continuous regression, but at the cost of inevitable approximation errors. Previous methods usually quantize the entire body pose into one code, which not only faces the difficulty in encoding all joints within one vector but a…

Cited by 2SourcePDFScholar
2024

Multi-View Representation is What You Need for Point-Cloud Pre-Training

ICLR 2024poster

A promising direction for pre-training 3D point clouds is to leverage the massive amount of data in 2D, whereas the domain gap between 2D and 3D creates a fundamental challenge. This paper proposes a novel approach to point-cloud pre-training that learns 3D representations by leveraging pre-trained…

Cited by 4SourcePDFScholar
2024

OmniGlue: Generalizable Feature Matching with Foundation Model Guidance

CVPR 2024poster

The image matching field has been witnessing a continuous emergence of novel learnable feature matching techniques with ever-improving performance on conventional benchmarks. However our investigation shows that despite these gains their potential for real-world applications is restricted by their l…

2024

PPLNs: Parametric Piecewise Linear Networks for Event-Based Temporal Modeling and Beyond

NeurIPS 2024poster

We present Parametric Piecewise Linear Networks (PPLNs) for temporal vision inference. Motivated by the neuromorphic principles that regulate biological neural behaviors, PPLNs are ideal for processing data captured by event cameras, which are built to simulate neural activities in the human retina.…

2024

TutteNet: Injective 3D Deformations by Composition of 2D Mesh Deformations

CVPR 2024highlight

This work proposes a novel representation of injective deformations of 3D space which overcomes existing limitations of injective methods namely inaccuracy lack of robustness and incompatibility with general learning and optimization frameworks. Our core idea is to reduce the problem to a "deep" com…

Cited by 0SourcePDFScholar
2024

ViGoR: Improving Visual Grounding of Large Vision Language Models with Fine-Grained Reward Modeling

ECCV 2024poster

"By combining natural language understanding, generation capabilities, and breadth of knowledge of large language models with image perception, recent large vision language models (LVLMs) have shown unprecedented visual reasoning capabilities. However, the generated text often suffers from inaccurat…

2023

Implicit Autoencoder for Point-Cloud Self-Supervised Representation Learning

ICCV 2023poster

This paper advocates the use of implicit surface representation in autoencoder-based self-supervised 3D representation learning. The most popular and accessible 3D representation, i.e., point clouds, involves discrete samples of the underlying continuous 3D surface. This discretization process intro…

Cited by 65PDFcodeScholar
2022

FvOR: Robust Joint Shape and Pose Optimization for Few-View Object Reconstruction

CVPR 2022poster

Reconstructing an accurate 3D object model from a few image observations remains a challenging problem in computer vision. State-of-the-art approaches typically assume accurate camera poses as input, which could be difficult to obtain in realistic settings. In this paper, we present FvOR, a learning…

Cited by 23PDFcodeScholar
2022

InfoGCN: Representation Learning for Human Skeleton-Based Action Recognition

CVPR 2022poster

Human skeleton-based action recognition offers a valuable means to understand the intricacies of human behavior because it can handle the complex relationships between physical constraints and intention. Although several studies have focused on encoding a skeleton, less attention has been paid to em…

Cited by 310PDFcodeScholar
2022

MVS2D: Efficient Multi-View Stereo via Attention-Driven 2D Convolutions

CVPR 2022poster

Deep learning has made significant impacts on multi-view stereo systems. State-of-the-art approaches typically involve building a cost volume, followed by multiple 3D convolution operations to recover the input image's pixel-wise depth. While such end-to-end learning of plane-sweeping stereo advance…

Cited by 59PDFcodeScholar
2022

PatchRD: Detail-Preserving Shape Completion by Learning Patch Retrieval and Deformation

ECCV 2022poster

"This paper introduces a data-driven shape completion approach that focuses on completing geometric details of missing regions of 3D shapes. We observe that existing generative methods do not have enough training data and representation capacity to synthesize plausible, fine-grained details with com…

2021

ARAPReg: An As-Rigid-As Possible Regularization Loss for Learning Deformable Shape Generators

ICCV 2021poster

This paper introduces an unsupervised loss for training parametric deformation shape generators. The key idea is to enforce the preservation of local rigidity among the generated shapes. Our approach builds on a local approximation of the as-rigid-as possible (or ARAP) deformation energy. We show ho…

Cited by 53PDFcodeScholar
2021

HPNet: Deep Primitive Segmentation Using Hybrid Representations

ICCV 2021poster

This paper introduces HPNet, a novel deep-learning approach for segmenting a 3D shape represented as a point cloud into primitive patches. The key to deep primitive segmentation is learning a feature representation that can separate points of different primitives. Unlike utilizing a single feature r…

Cited by 56PDFcodeScholar
2021

Scene Synthesis via Uncertainty-Driven Attribute Synchronization

ICCV 2021poster

Developing deep neural networks to generate 3D scenes is a fundamental problem in neural synthesis with immediate applications in architectural CAD, computer graphics, as well as in generating virtual robot training environments. This task is challenging because 3D scenes exhibit diverse patterns, r…

Cited by 39PDFcodeScholar
2020

A Large-scale Annotated Mechanical Components Benchmark for Classification and Retrieval Tasks with Deep Neural Networks

ECCV 2020poster

We introduce a large-scale annotated mechanical components benchmark for classification and retrieval tasks named MechanicalComponents Benchmark (MCB): a large-scale dataset of 3D objects of mechanical components. The dataset enables data-driven feature learn-ing for mechanical components. Exploring…

2020

Dense Correspondences between Human Bodies via Learning Transformation Synchronization on Graphs

NeurIPS 2020poster

We introduce an approach for establishing dense correspondences between partial scans of human models and a complete template model. Our approach's key novelty lies in formulating dense correspondence computation as initializing and synchronizing local transformations between the scan and the templa…

2020

H3DNet: 3D Object Detection Using Hybrid Geometric Primitives

ECCV 2020poster

We introduce H3DNet, which takes a colorless 3D point cloud as input and outputs a collection of oriented object bounding boxes (or BB) and their semantic labels. The critical idea of H3DNet is to predict a hybrid set of geometric primitives, i.e., BB centers, BB face centers, and BB edge centers. W…

2019

A Condition Number for Joint Optimization of Cycle-Consistent Networks

NeurIPS 2019spotlight

A recent trend in optimizing maps such as dense correspondences between objects or neural networks between pairs of domains is to optimize them jointly. In this context, there is a natural \textsl{cycle-consistency} constraint, which regularizes composite maps associated with cycles, i.e., they are…

2019

Extreme Relative Pose Estimation for RGB-D Scans via Scene Completion

CVPR 2019oral

Estimating the relative rigid pose between two RGB-D scans of the same underlying environment is a fundamental problem in computer vision, robotics, and computer graphics. Most existing approaches allow only limited maximum relative pose changes since they require considerable overlap between the in…

Cited by 57PDFcodeScholar
2019

Fast and Robust Multi-Person 3D Pose Estimation From Multiple Views

CVPR 2019poster

This paper addresses the problem of 3D pose estimation for multiple people in a few calibrated camera views. The main challenge of this problem is to find the cross-view correspondences among noisy and incomplete 2D pose predictions. Most previous methods address this challenge by directly reasoning…

Cited by 266PDFScholar
2019

Learning Transformation Synchronization

CVPR 2019poster

Reconstructing the 3D model of a physical object typically requires us to align the depth scans obtained from different camera poses into the same coordinate system. Solutions to this global alignment problem usually proceed in two steps. The first step estimates relative transformations between pai…

Cited by 67PDFcodeScholar
2019

PVNet: Pixel-Wise Voting Network for 6DoF Pose Estimation

CVPR 2019oral

This paper addresses the challenge of 6DoF pose estimation from a single RGB image under severe occlusion or truncation. Many recent works have shown that a two-stage approach, which first detects keypoints and then solves a Perspective-n-Point (PnP) problem for pose estimation, achieves remarkable…

Cited by 1355PDFcodeScholar
2018

StarMap for Category-Agnostic Keypoint and Viewpoint Estimation

ECCV 2018poster

Semantic keypoints provide concise abstractions for a variety of visual understanding tasks. Existing methods define semantic keypoints separately for each category with a fixed number of semantic labels in fixed indices. As a result, this keypoint representation is in-feasible when objects have a v…

2018

Unsupervised Domain Adaptation for 3D Keypoint Estimation via View Consistency

ECCV 2018poster

In this paper, we introduce a novel unsupervised domain adaptation technique for the task of 3D keypoint prediction from a single depth scan or image. Our key idea is to utilize the fact that predictions from different views of the same or similar objects should be consistent with each other. Such v…

2017

Greedy Direction Method of Multiplier for MAP Inference of Large Output Domain

AISTATS 2017poster

Maximum-a-Posteriori (MAP) inference lies at the heart of Graphical Models and Structured Prediction. Despite the intractability of exact MAP inference, approximated methods based on LP relaxations have exhibited superior performance across a wide range of applications. Yet for problems involving la…

Cited by 7SourcePDFScholar
2017

SurfNet: Generating 3D Shape Surfaces Using Deep Residual Networks

CVPR 2017poster

3D shape models are naturally parameterized using vertices and faces, i.e, composed on polygons forming a surface. However, current 3D learning paradigms for predictive and generative tasks using convolutional neural networks focus on a voxelized representation of the object. Lifting convolution ope…

Cited by 213PDFcodeScholar
2017

Towards 3D Human Pose Estimation in the Wild: A Weakly-Supervised Approach

ICCV 2017poster

In this paper, we study the task of 3D human pose estimation in the wild. This task is challenging due to lack of training data, as existing datasets are either in the wild images with 2D pose or in the lab images with 3D pose. We propose a weakly-supervised transfer learning method that uses mixed…

Cited by 749PDFcodeScholar
2017

Translation Synchronization via Truncated Least Squares

NeurIPS 2017spotlight

In this paper, we introduce a robust algorithm, \textsl{TranSync}, for the 1D translation synchronization problem, in which the aim is to recover the global coordinates of a set of nodes from noisy measurements of relative coordinates along an observation graph. The basic idea of TranSync is to appl…

Cited by 45SourcePDFScholar
2016

Dense Human Body Correspondences Using Convolutional Networks

CVPR 2016oral

We propose a deep learning approach for finding dense correspondences between 3D scans of people. Our method requires only partial geometric information in the form of two depth maps or partial reconstructed surfaces, works for humans in arbitrary poses and wearing any clothing, does not require the…

Cited by 242PDFScholar
2016

Learning Dense Correspondence via 3D-Guided Cycle Consistency

CVPR 2016oral

Discriminative deep learning approaches have shown impressive results for problems where human-labeled ground truth is plentiful, but what about tasks where labels are difficult or impossible to obtain? This paper tackles one such problem: establishing dense visual correspondence across different ob…

Cited by 453PDFScholar