← Search

Xiaoming Zhao

14 accepted papers

2026

Velox: Learning Representations of 4D Geometry and Appearance

CVPR 2026

We introduce a framework for learning latent representations of 4D objects which are descriptive, faithfully capturing object geometry and appearance; compressive, aiding in downstream efficiency; and accessible, requiring minimal input, i.e., an unstructured dynamic point cloud, to construct. Speci

Cited by 0SourceScholar
2025

PDC & DM-SFT: A Road for LLM SQL Bug-Fix Enhancing

COLING 2025industry

Code Large Language Models (Code LLMs), such as Code llama and DeepSeek-Coder, have demonstrated exceptional performance in the code generation tasks. However, most existing models focus on the abilities of generating correct code, but often struggle with bug repair. We introduce a suit of methods t…

2024

GoMAvatar: Efficient Animatable Human Modeling from Monocular Video Using Gaussians-on-Mesh

CVPR 2024poster

We introduce GoMAvatar a novel approach for real-time memory-efficient high-quality animatable human modeling. GoMAvatar takes as input a single monocular video to create a digital avatar capable of re-articulation in new poses and real-time rendering from novel viewpoints while seamlessly integrati…

Cited by 34SourcePDFScholar
2024

IllumiNeRF: 3D Relighting Without Inverse Rendering

NeurIPS 2024poster

Existing methods for relightable view synthesis --- using a set of images of an object under unknown lighting to recover a 3D representation that can be rendered from novel viewpoints under a target illumination --- are based on inverse rendering, and attempt to disentangle the object geometry, mate…

2024

Integrating Representation Subspace Mapping with Unimodal Auxiliary Loss for Attention-based Multimodal Emotion Recognition

COLING 2024main

Multimodal emotion recognition (MER) aims to identify emotions by utilizing affective information from multiple modalities. Due to the inherent disparities among these heterogeneous modalities, there is a large modality gap in their representations, leading to the challenge of fusing multiple modali…

Cited by 1SourcePDFScholar
2024

NeRFDeformer: NeRF Transformation from a Single View via 3D Scene Flows

CVPR 2024poster

We present a method for automatically modifying a NeRF representation based on a single observation of a non-rigid transformed version of the original scene. Our method defines the transformation as a 3D flowspecifically as a weighted linear blending of rigid transformations of 3D anchor points that…

2024

Pseudo-Generalized Dynamic View Synthesis from a Video

ICLR 2024poster

Rendering scenes observed in a monocular video from novel viewpoints is a challenging problem. For static scenes the community has studied both scene-specific optimization techniques, which optimize on every test scene, and generalized techniques, which only run a deep net forward pass on a test sce…

2023

Occupancy Planes for Single-View RGB-D Human Reconstruction

AAAI 2023technical

Single-view RGB-D human reconstruction with implicit functions is often formulated as per-point classification. Specifically, a set of 3D locations within the view-frustum of the camera are first projected independently onto the image and a corresponding feature is subsequently extracted for each 3…

2022

Generative Multiplane Images: Making a 2D GAN 3D-Aware

ECCV 2022poster

"What is really needed to make an existing 2D GAN 3Daware? To answer this question, we modify a classical GAN, i.e., StyleGANv2, as little as possible. We find that only two modifications are absolutely necessary: 1) a multiplane image style generator branch which produces a set of alpha maps condit…

2022

Initialization and Alignment for Adversarial Texture Optimization

ECCV 2022poster

"While recovery of geometry from image and video data has received a lot of attention in computer vision, methods to capture the texture for a given geometry are less mature. Specifically, classical methods for texture generation often assume clean geometry and reasonably well-aligned image data. Wh…

2021

The Surprising Effectiveness of Visual Odometry Techniques for Embodied PointGoal Navigation

ICCV 2021poster

It is fundamental for personal robots to reliably navigate to a specified goal. To study this task, PointGoal navigation has been introduced in simulated Embodied AI environments. Recent advances solve this PointGoal navigation task with near-perfect accuracy (99.6% success) in photo-realistically s…

Cited by 53PDFScholar