← Search

Xun Cao

55 accepted papers

2026

ComGS: Efficient 3D Object-Scene Composition via Surface Octahedral Probes

ICLR 2026poster

Gaussian Splatting (GS) enables immersive rendering, but realistic 3D object–scene composition remains challenging. Baked appearance and shadow information in GS radiance fields cause inconsistencies when combining objects and scenes. Addressing this requires relightable object reconstruction and sc…

Cited by 0SourcecodeScholar
2026

DeX-Portrait: Disentangled and Expressive Portrait Animation via Explicit and Latent Motion Representations

CVPR 2026

Portrait animation from a single source image and a driving video is a long-standing problem. Recent approaches tend to adopt diffusion-based image/video generation models for realistic and expressive animation. However, none of these diffusion models realizes high-fidelity disentangled control betw

Cited by 0SourceScholar
2026

Exploring Spatiotemporal Feature Propagation for Video-Level Compressive Spectral Reconstruction: Dataset, Model and Benchmark

CVPR 2026

Recently, Spectral Compressive Imaging (SCI) has achieved remarkable success, unlocking significant potential for dynamic spectral vision. However, existing reconstruction methods, primarily image-based, suffer from two limitations: (i) Encoding process masks spatial-spectral features, leading to un

Cited by 0SourcecodeScholar
2026

FastGHA: Generalized Few-Shot 3D Gaussian Head Avatars with Real-Time Animation

ICLR 2026poster

Despite recent progress in 3D Gaussian-based head avatar modeling, efficiently generating high fidelity avatars remains a challenge. Current methods typically rely on extensive multi-view capture setups or monocular videos with per-identity optimization during inference, limiting their scalability a…

Cited by 0SourceScholar
2026

LiteVGGT: Boosting Vanilla VGGT via Geometry-aware Cached Token Merging

CVPR 2026

3D vision foundation models like Visual Geometry Grounded Transformer (VGGT) have advanced greatly in geometric perception. However it is time-consuming and memory-intensive for long sequences, limiting application to large-scale scenes beyond hundreds of images. To address this, we propose LiteVGGT

Cited by 0SourcecodeScholar
2026

Pressure2Motion: Hierarchical Human Motion Reconstruction from Ground Pressure with Text Guidance

CVPR 2026

We present Pressure2Motion, a novel motion capture algorithm that reconstructs human motion from a ground pressure sequence and text prompt. At inference time, Pressure2Motion requires only a pressure mat, eliminating the need for specialized lighting setups, cameras, or wearable devices, making it

Cited by 0SourcecodeScholar
2026

SpatialVID: A Large-Scale Video Dataset with Spatial Annotations

CVPR 2026

Significant progress has been made in spatial intelligence, spanning both spatial reconstruction and world exploration. However, the scalability and real-world fidelity of current models remain severely constrained by the scarcity of large-scale, high-quality training data. While several datasets pr

Cited by 0SourcecodeScholar
2026

Spike Imaging Velocimetry: Dense Motion Estimation of Fluids Using Spike Streams

AAAI 2026technical

Particle Image Velocimetry (PIV) is a widely adopted non-invasive imaging technique that tracks the motion of tracer particles across image sequences to capture the velocity distribution of fluid flows. It is commonly employed to analyze complex flow structures and validate numerical simulations. Th

Cited by 0SourcePDFScholar
2026

Split-Layer: Enhancing Implicit Neural Representation by Maximizing the Dimensionality of Feature Space

AAAI 2026technical

Implicit neural representation (INR) models signals as continuous functions using neural networks, offering efficient and differentiable optimization for inverse problems across diverse disciplines. However, the representational capacity of INR—defined by the range of functions the neural network ca

Cited by 0SourcePDFScholar
2026

TEXTRIX: Latent Attribute Grid for Native Texture Generation and Beyond

CVPR 2026

Prevailing 3D texture generation methods, which often rely on multi-view fusion, are frequently hindered by inter-view inconsistencies and incomplete coverage of complex surfaces, limiting the fidelity and completeness of the generated content. To overcome these challenges, we introduce TEXTRIX, a n

Cited by 0SourcecodeScholar
2026

UIKA: Fast Universal Head Avatar from Pose-Free Images

CVPR 2026

We present UIKA, a feed-forward animatable Gaussian head model from an arbitrary number of pose-free inputs, including a single image, multi-view captures, and smartphone-captured videos. Unlike the traditional avatar method, which requires a studio-level multi-view capture system and reconstructs a

Cited by 0SourcecodeScholar
2025

Direct3D-S2: Gigascale 3D Generation Made Easy with Spatial Sparse Attention

NeurIPS 2025poster

Generating high-resolution 3D shapes using volumetric representations such as Signed Distance Functions (SDFs) presents substantial computational and memory challenges. We introduce Direct3D-S2, a scalable 3D generation framework based on sparse volumes that achieves superior output quality with dra…

Cited by 0SourceScholar
2025

Epona: Autoregressive Diffusion World Model for Autonomous Driving

ICCV 2025poster

Diffusion models have demonstrated exceptional visual quality in video generation, making them promising for autonomous driving world modeling. However, existing video diffusion-based world models struggle with flexible-length, long-horizon predictions and integrating trajectory planning. This is be…

2025

FATE: Full-head Gaussian Avatar with Textural Editing from Monocular Video

CVPR 2025poster

Reconstructing high-fidelity, animatable 3D head avatars from effortlessly captured monocular videos is a pivotal yet formidable challenge. Although significant progress has been made in rendering performance and manipulation capabilities, notable challenges remain, including incomplete reconstructi…

2025

Flow Distillation Sampling: Regularizing 3D Gaussians with Pre-trained Matching Priors

ICLR 2025poster

3D Gaussian Splatting (3DGS) has achieved excellent rendering quality with fast training and rendering speed. However, its optimization process lacks explicit geometric constraints, leading to suboptimal geometric reconstruction in regions with sparse or no observational input views. In this work, w…

Cited by 0SourcePDFScholar
2025

IDOL: Instant Photorealistic 3D Human Creation from a Single Image

CVPR 2025poster

Creating a high-fidelity, animatable 3D full-body avatar from a single image is a challenging task due to the diverse appearance and poses of humans and the limited availability of high-quality training data. To achieve fast and high-quality human reconstruction, this work rethinks the task from the…

2025

M-SpecGene: Generalized Foundation Model for RGBT Multispectral Vision

ICCV 2025poster

RGB-Thermal (RGBT) multispectral vision is essential for robust perception in complex environments. Most RGBT tasks follow a case-by-case research paradigm, relying on manually customized models to learn task-oriented representations. Nevertheless, this paradigm is inherently constrained by artifici…

Cited by 0SourcePDFScholar
2025

Matrix3D: Large Photogrammetry Model All-in-One

CVPR 2025highlight

We present Matrix3D, a unified model that performs several photogrammetry subtasks, including pose estimation, depth prediction, and novel view synthesis using just the same model. Matrix3D utilizes a multi-modal diffusion transformer (DiT) to integrate transformations across several modalities, suc…

2025

Mitigating Ambiguities in 3D Classification with Gaussian Splatting

CVPR 2025poster

3D classification with point cloud input is a fundamental problem in 3D vision. However, due to the discrete nature and the insufficient material description of point cloud representations, there are ambiguities in distinguishing wire-like and flat surfaces, as well as transparent or reflective obje…

Cited by 0SourcePDFScholar
2025

MotionPRO: Exploring the Role of Pressure in Human MoCap and Beyond

CVPR 2025highlight

Existing human Motion Capture (MoCap) methods mostly focus on the visual similarity while neglecting the physical plausibility. As a result, downstream tasks such as driving virtual human in 3D scene or humanoid robots in real world suffer from issues such as timing drift and jitter, spatial problem…

2025

TeRA: Rethinking Text-guided Realistic 3D Avatar Generation

ICCV 2025poster

Efficient 3D avatar creation is a significant demand in the metaverse, film/game, AR/VR, etc. In this paper, we rethink text-to-avatar generative models by proposing TeRA, a more efficient and effective framework than the previous SDS-based models and general large 3D generative models. Our approach…

Cited by 0SourcePDFScholar
2024

A Pre-convolved Representation for Plug-and-Play Neural Illumination Fields

AAAI 2024technical

Recent advances in implicit neural representation have demonstrated the ability to recover detailed geometry and material from multi-view images. However, the use of simplified lighting models such as environment maps to represent non-distant illumination, or using a network to fit indirect light mo…

Cited by 2SourcePDFScholar
2024

Batch Normalization Alleviates the Spectral Bias in Coordinate Networks

CVPR 2024poster

Representing signals using coordinate networks dominates the area of inverse problems recently and is widely applied in various scientific computing tasks. Still there exists an issue of spectral bias in coordinate networks limiting the capacity to learn high-frequency components. This problem is ca…

Cited by 9SourcePDFScholar
2024

Champ: Controllable and Consistent Human Image Animation with 3D Parametric Guidance

ECCV 2024poster

"In this study, we introduce a methodology for human image animation by leveraging a 3D human parametric model within a latent diffusion framework to enhance shape alignment and motion guidance in current human generative techniques. The methodology utilizes the SMPL(Skinned Multi-Person Linear) mod…

2024

Direct2.5: Diverse Text-to-3D Generation via Multi-view 2.5D Diffusion

CVPR 2024poster

Recent advances in generative AI have unveiled significant potential for the creation of 3D content. However current methods either apply a pre-trained 2D diffusion model with the time-consuming score distillation sampling (SDS) or a direct 3D diffusion model trained on limited 3D data losing genera…

Cited by 33SourcePDFScholar
2024

Direct3D: Scalable Image-to-3D Generation via 3D Latent Diffusion Transformer

NeurIPS 2024poster

Generating high-quality 3D assets from text and images has long been challenging, primarily due to the absence of scalable 3D representations capable of capturing intricate geometry distributions. In this work, we introduce Direct3D, a native 3D generative model scalable to in-the-wild input images,…

Cited by 35SourcePDFScholar
2024

Efficient Snapshot Spectral Imaging: Calibration-Free Parallel Structure with Aperture Diffraction Fusion

ECCV 2024poster

"Aiming to address the repetitive need for complex calibration in existing snapshot spectral imaging methods, while better trading off system complexity and spectral reconstruction accuracy, we demonstrate a novel Parallel Coded Calibration-free Aperture Diffraction Imaging Spectrometer (PCCADIS) wi…

Cited by 0SourcePDFScholar
2024

EmoTalk3D: High-Fidelity Free-View Synthesis of Emotional 3D Talking Head

ECCV 2024poster

"We present a novel approach for synthesizing 3D talking heads with controllable emotion, featuring enhanced lip synchronization and rendering quality. Despite significant progress in the field, prior methods still suffer from multi-view consistency and a lack of emotional expressiveness. To address…

2024

FINER: Flexible Spectral-bias Tuning in Implicit NEural Representation by Variable-periodic Activation Functions

CVPR 2024poster

Implicit Neural Representation (INR) which utilizes a neural network to map coordinate inputs to corresponding attributes is causing a revolution in the field of signal processing. However current INR techniques suffer from a restricted capability to tune their supported frequency set resulting in i…

Cited by 34SourcePDFScholar
2024

Head360: Learning a Parametric 3D Full-Head for Free-View Synthesis in 360°

ECCV 2024poster

"Creating a 360◦ parametric model of a human head is a very challenging task. While recent advancements have demonstrated the efficacy of leveraging synthetic data for building such parametric head models, their performance remains inadequate in crucial areas such as expression-driven animation, hai…

2024

Joint RGB-Spectral Decomposition Model Guided Image Enhancement in Mobile Photography

ECCV 2024poster

"The integration of miniaturized spectrometers into mobile devices offers new avenues for image quality enhancement and facilitates novel downstream tasks. However, the broader application of spectral sensors in mobile photography is hindered by the inherent complexity of spectral images and the con…

2024

MMVP: A Multimodal MoCap Dataset with Vision and Pressure Sensors

CVPR 2024poster

Foot contact is an important cue for human motion capture understanding and generation. Existing datasets tend to annotate dense foot contact using visual matching with thresholding or incorporating pressure signals. However these approaches either suffer from low accuracy or are only designed for s…

2024

NeRI: Implicit Neural Representation of LiDAR Point Cloud Using Range Image Sequence

ICASSP 2024accepted

This paper proposes the NeRI, an implicit neural representation (INR) based LiDAR point cloud compressor. In NeRI, we first transform a sequence of 3D LiDAR frames into a 2D range image sequence through range image projection over time. Then, we employ a neural network conditioned on the temporal fr…

Cited by 0SourceScholar
2024

Relightable 3D Gaussians: Realistic Point Cloud Relighting with BRDF Decomposition and Ray Tracing

ECCV 2024poster

"In this paper, we present a novel differentiable point-based rendering framework to achieve photo-realistic relighting. To make the reconstructed scene relightable, we enhance vanilla 3D Gaussians by associating extra properties, including normal vectors, BRDF parameters, and incident lighting from…

Cited by 133SourcePDFScholar
2024

STAG4D: Spatial-Temporal Anchored Generative 4D Gaussians

ECCV 2024poster

"Recent progress in pre-trained diffusion models and 3D generation have spurred interest in 4D content creation. However, achieving high-fidelity 4D generation with spatial-temporal consistency remains a challenge. In this work, we propose STAG4D, a novel framework that combines pre-trained diffusio…

Cited by 46SourcePDFScholar
2023

Aperture Diffraction for Compact Snapshot Spectral Imaging

ICCV 2023poster

We demonstrate a compact, cost-effective snapshot spectral imaging system named Aperture Diffraction Imaging Spectrometer (ADIS), which consists only of an imaging lens with an ultra-thin orthogonal aperture mask and a mosaic filter sensor, requiring no additional physical footprint compared to comm…

Cited by 4PDFcodeScholar
2023

DINER: Disorder-Invariant Implicit Neural Representation

CVPR 2023highlight

Implicit neural representation (INR) characterizes the attributes of a signal as a function of corresponding coordinates which emerges as a sharp weapon for solving inverse problems. However, the capacity of INR is limited by the spectral bias in the network training. In this paper, we find that suc…

2023

High-Fidelity 3D Face Generation From Natural Language Descriptions

CVPR 2023poster

Synthesizing high-quality 3D face models from natural language descriptions is very valuable for many applications, including avatar creation, virtual reality, and telepresence. However, little research ever tapped into this task. We argue the major obstacle lies in 1) the lack of high-quality 3D fa…

2023

RAFaRe: Learning Robust and Accurate Non-parametric 3D Face Reconstruction from Pseudo 2D&3D Pairs

AAAI 2023technical

We propose a robust and accurate non-parametric method for single-view 3D face reconstruction (SVFR). While tremendous efforts have been devoted to parametric SVFR, a visible gap still lies between the result 3D shape and the ground truth. We believe there are two major obstacles: 1) the representat…

2022

Detailed Facial Geometry Recovery from Multi-View Images by Learning an Implicit Function

AAAI 2022technical

Recovering detailed facial geometry from a set of calibrated multi-view images is valuable for its wide range of applications. Traditional multi-view stereo (MVS) methods adopt an optimization-based scheme to regularize the matching cost. Recently, learning-based methods integrate all these into an…

2022

Explore Spatio-Temporal Aggregation for Insubstantial Object Detection: Benchmark Dataset and Baseline

CVPR 2022poster

We endeavor on a rarely explored task named Insubstan-tial Object Detection (IOD), which aims to localize the object with following characteristics: (1) amorphous shape with indistinct boundary; (2) similarity to surroundings; (3) absence in color. Accordingly, it is far more challenging to distingu…

Cited by 23PDFcodeScholar
2020

FaceScape: A Large-Scale High Quality 3D Face Dataset and Detailed Riggable 3D Face Prediction

CVPR 2020poster

In this paper, we present a large-scale detailed 3D face dataset, FaceScape, and propose a novel algorithm that is able to predict elaborate riggable 3D face models from a single image input. FaceScape dataset provides 18,760 textured 3D faces, captured from 938 subjects and each with 20 specific ex…

Cited by 361PDFcodeScholar
2020

Improving Multispectral Pedestrian Detection by Addressing Modality Imbalance Problems

ECCV 2020poster

Multispectral pedestrian detection is capable of adapting to insufficient illumination conditions by leveraging color-thermal modalities. On the other hand, it is still lacking of in-depth insights on how to fuse the two modalities effectively. Compared with traditional pedestrian detection, we find…

2019

Detailed Human Shape Estimation From a Single Image by Hierarchical Mesh Deformation

CVPR 2019oral

This paper presents a novel framework to recover detailed human body shapes from a single image. It is a challenging task due to factors such as variations in human shapes, body poses, and viewpoints. Prior methods typically attempt to recover the human body shape using a parametric based template…

Cited by 173PDFcodeScholar
2019

Spectral Reconstruction From Dispersive Blur: A Novel Light Efficient Spectral Imager

CVPR 2019poster

Developing high light efficiency imaging techniques to retrieve high dimensional optical signal is a long-term goal in computational photography. Multispectral imaging, which captures images of different wavelengths and boosting the abilities for revealing scene properties, has developed rapidly in…

Cited by 5PDFScholar
2018

Multispectral Image Intrinsic Decomposition via Subspace Constraint

CVPR 2018poster

Multispectral images contain many clues of surface characteristics of the objects, thus can be used in many computer vision tasks, e.g., recolorization and segmentation. However, due to the complex geometry structure of natural scenes, the spectra curves of the same surface can look very different u…

Cited by 13SourcePDFScholar
2015

Blind Optical Aberration Correction by Exploring Geometric and Visual Priors

CVPR 2015poster

Optical aberration widely exists in optical imaging systems, especially in consumer-level cameras. In contrast to previous solutions using hardware compensation or pre-calibration, we propose a computational approach for blind aberration removal from a single image, by exploring various geometric an…

Cited by 45SourcePDFScholar