← Search

Yinda Zhang

51 accepted papers

2026

Archon: A Unified Multimodal Model for Holistic Digital Human Generation

CVPR 2026

Digital humans are fundamental to immersive interaction, yet creating a unified model for holistic modalities, including text, audio, motion, and visual content, remains an open challenge. In this paper, we present Archon, a fully pretrained, human-centric unified multimodal model for holistic avata

Cited by 0SourceScholar
2026

Talking Together: Synthesizing Co-Located 3D Conversations from Audio

CVPR 2026

We tackle the challenging task of generating complete 3D facial animations for two interacting, co-located participants from a mixed audio stream. While existing methods often produce disembodied "talking heads" akin to a video conference call, our work is the first to explicitly model the dynamic 3

Cited by 0SourceScholar
2025

DiffGrasp: Whole-Body Grasping Synthesis Guided by Object Motion Using a Diffusion Model

AAAI 2025technical

Generating high-quality whole-body human object interaction motion sequences is becoming increasingly important in various fields such as animation, VR/AR, and robotics. The main challenge of this task lies in determining the level of involvement of each hand given the complex shapes of objects in d…

Cited by 1SourcePDFScholar
2025

EVER: Exact Volumetric Ellipsoid Rendering for Real-time View Synthesis

ICCV 2025poster

We present Exact Volumetric Ellipsoid Rendering (EVER), a method for real-time 3D reconstruction.EVER accurately blends an unlimited number of overlapping primitives together in 3D space, eliminating the popping artifacts that 3D Gaussian Splatting (3DGS) and other related methods exhibit.EVER repre…

Cited by 0SourcePDFScholar
2025

HOGSA: Bimanual Hand-Object Interaction Understanding with 3D Gaussian Splatting Based Data Augmentation

AAAI 2025technical

Understanding of bimanual hand-object interaction plays an important role in robotics and virtual reality. However, due to significant occlusions between hands and object as well as the high degree-of-freedom motions, it is challenging to collect and annotate a high-quality, large-scale dataset, whi…

Cited by 0SourcePDFScholar
2025

IM-Portrait: Learning 3D-aware Video Diffusion for Photorealistic Talking Heads from Monocular VideosC

CVPR 2025poster

We propose a novel 3D-aware diffusion-based method for generating photorealistic talking head videos directly from a single identity image and explicit control signals (e.g., expressions). Our method generates Multiplane Images (MPIs) that ensure geometric consistency, making them ideal for immersiv…

Cited by 0SourcePDFScholar
2025

SVG: 3D Stereoscopic Video Generation via Denoising Frame Matrix

ICLR 2025poster

Video generation models have demonstrated great capability of producing impressive monocular videos, however, the generation of 3D stereoscopic video remains under-explored. We propose a pose-free and training-free approach for generating 3D stereoscopic videos using an off-the-shelf monocular video…

2024

Efficient 3D Implicit Head Avatar with Mesh-anchored Hash Table Blendshapes

CVPR 2024poster

3D head avatars built with neural implicit volumetric representations have achieved unprecedented levels of photorealism. However the computational cost of these methods remains a significant barrier to their widespread adoption particularly in real-time applications such as virtual reality and tele…

Cited by 5SourcePDFScholar
2024

GeneAvatar: Generic Expression-Aware Volumetric Head Avatar Editing from a Single Image

CVPR 2024poster

Recently we have witnessed the explosive growth of various volumetric representations in modeling animatable head avatars. However due to the diversity of frameworks there is no practical method to support high-level applications like 3D head avatar editing across different representations. In this…

2024

MVDD: Multi-View Depth Diffusion Models

ECCV 2024poster

"Denoising diffusion models have demonstrated outstanding results in 2D image generation, yet it remains a challenge to replicate its success in 3D shape generation. In this paper, we propose leveraging multi-view depth, which represents complex 3D shapes in a 2D data format that is easy to denoise.…

Cited by 4SourcePDFScholar
2024

T-Pixel2Mesh: Combining Global and Local Transformer for 3D Mesh Generation from a Single Image

ICASSP 2024accepted

Pixel2Mesh (P2M) is a classical approach for reconstructing 3D shapes from a single color image through coarse-to-fine mesh deformation. Although P2M is capable of generating plausible global shapes, its Graph Convolution Network (GCN) often produces overly smooth results, causing the loss of fine-g…

Cited by 0SourceScholar
2023

Grad-PU: Arbitrary-Scale Point Cloud Upsampling via Gradient Descent With Learned Distance Functions

CVPR 2023poster

Most existing point cloud upsampling methods have roughly three steps: feature extraction, feature expansion and 3D coordinate prediction. However, they usually suffer from two critical issues: (1) fixed upsampling rate after one-time training, since the feature expansion unit is customized for each…

2023

Hybrid Neural Rendering for Large-Scale Scenes With Motion Blur

CVPR 2023poster

Rendering novel view images is highly desirable for many applications. Despite recent progress, it remains challenging to render high-fidelity and view-consistent novel views of large-scale scenes from in-the-wild images with inevitable artifacts (e.g., motion blur). To this end, we develop a hybrid…

2023

Learning Personalized High Quality Volumetric Head Avatars From Monocular RGB Videos

CVPR 2023poster

We propose a method to learn a high-quality implicit 3D head avatar from a monocular RGB video captured in the wild. The learnt avatar is driven by a parametric face model to achieve user-controlled facial expressions and head poses. Our hybrid pipeline combines the geometry prior and dynamic tracki…

Cited by 20SourcePDFScholar
2023

Learning Versatile 3D Shape Generation with Improved Auto-regressive Models

ICCV 2023poster

Auto-Regressive (AR) models have achieved impressive results in 2D image generation by modeling joint distributions in the grid space. While this approach has been extended to the 3D domain for powerful shape generation, it still has two limitations: expensive computations on volumetric grids and am…

Cited by 1PDFScholar
2023

Multi-Modal Neural Radiance Field for Monocular Dense SLAM with a Light-Weight ToF Sensor

ICCV 2023poster

Light-weight time-of-flight (ToF) depth sensors are compact and cost-efficient, and thus widely used on mobile devices for tasks such as autofocus and obstacle detection. However, due to the sparse and noisy depth measurements, these sensors have rarely been considered for dense geometry reconstruct…

Cited by 32PDFcodeScholar
2023

Novel-View Synthesis and Pose Estimation for Hand-Object Interaction from Sparse Views

ICCV 2023poster

Hand-object interaction understanding and the barely addressed novel view synthesis are highly desired in the immersive communication, whereas it is challenging due to the high deformation of hand and heavy occlusions between hand and object. In this paper, we propose a neural rendering and pose est…

Cited by 16PDFcodeScholar
2023

SINE: Semantic-Driven Image-Based NeRF Editing With Prior-Guided Editing Field

CVPR 2023poster

Despite the great success in 2D editing using user-friendly tools, such as Photoshop, semantic strokes, or even text prompts, similar capabilities in 3D areas are still limited, either relying on 3D modeling skills or allowing editing within only a few categories. In this paper, we present a novel s…

2023

Self-supervised Learning of Implicit Shape Representation with Dense Correspondence for Deformable Objects

ICCV 2023poster

Learning 3D shape representation with dense correspondence for deformable objects is a fundamental problem in computer vision. Existing approaches often need additional annotations of specific semantic domain, e.g., skeleton pose for human body or animals, which require extra annotation effort and s…

Cited by 9PDFScholar
2023

Spectral Graphormer: Spectral Graph-Based Transformer for Egocentric Two-Hand Reconstruction using Multi-View Color Images

ICCV 2023poster

We propose a novel transformer-based framework that reconstructs two high fidelity hands from multi-view RGB images. Unlike existing hand pose estimation methods, where one typically trains a deep network to regress hand model parameters from single RGB image, we consider a more challenging problem…

Cited by 4PDFScholar
2022

DELTAR: Depth Estimation from a Light-Weight ToF Sensor and RGB Image

ECCV 2022poster

"Light-weight time-of-flight (ToF) depth sensors are small, cheap, low-energy and have been massively deployed on mobile devices for the purposes like autofocus, obstacle detection, etc. However, due to their specific measurements (depth distribution in a region instead of the depth value at a certa…

2022

Density-Preserving Deep Point Cloud Compression

CVPR 2022poster

Local density of point clouds is crucial for representing local details, but has been overlooked by existing point cloud compression methods. To address this, we propose a novel deep point cloud compression method that preserves local density information. Our method works in an auto-encoder fashion:…

Cited by 70PDFcodeScholar
2022

Efficient Virtual View Selection for 3D Hand Pose Estimation

AAAI 2022technical

3D hand pose estimation from single depth is a fundamental problem in computer vision, and has wide applications. However, the existing methods still can not achieve satisfactory hand pose estimation results due to view variation and occlusion of human hand. In this paper, we propose a new virtual v…

2022

H4D: Human 4D Modeling by Learning Neural Compositional Representation

CVPR 2022poster

Despite the impressive results achieved by deep learning based 3D reconstruction, the techniques of directly learning to model 4D human captures with detailed geometry have been less studied. This work presents a novel framework that can effectively learn a compact and compositional representation f…

Cited by 26PDFScholar
2022

LoRD: Local 4D Implicit Representation for High-Fidelity Dynamic Human Modeling

ECCV 2022poster

"Recent progress in 4D implicit representation focuses on globally controlling the shape and motion with low dimensional latent vectors, which is prone to missing surface details and accumulating tracking error. While many deep local representations have shown promising results for 3D shape modeling…

2022

NeuMesh: Learning Disentangled Neural Mesh-Based Implicit Field for Geometry and Texture Editing

ECCV 2022poster

"Very recently neural implicit rendering techniques have been rapidly evolved and shown great advantages in novel view synthesis and 3D scene reconstruction. However, existing neural rendering methods for editing purposes offer limited functionality, e.g., rigid transformation, or not applicable for…

2022

PRIF: Primary Ray-Based Implicit Function

ECCV 2022poster

"We introduce a new implicit shape representation called Primary Ray-based Implicit Function (PRIF). In contrast to most existing approaches based on the signed distance function (SDF) which handles spatial locations, our representation operates on oriented rays. Specifically, PRIF is formulated to…

Cited by 56SourcePDFScholar
2021

DeepPanoContext: Panoramic 3D Scene Understanding With Holistic Scene Context Graph and Relation-Based Optimization

ICCV 2021poster

Panorama images have a much larger field-of-view thus naturally encode enriched scene context information compared to standard perspective images, which however is not well exploited in the previous scene understanding methods. In this paper, we propose a novel method for panoramic 3D scene understa…

Cited by 40PDFcodeScholar
2021

HITNet: Hierarchical Iterative Tile Refinement Network for Real-time Stereo Matching

CVPR 2021poster

This paper presents HITNet, a novel neural network architecture for real-time stereo matching. Contrary to many recent neural network approaches that operate on a full costvolume and rely on 3D convolutions, our approach does not explicitly build a volume and instead relies on a fast multi-resolutio…

Cited by 326PDFcodeScholar
2021

Holistic 3D Scene Understanding From a Single Image With Implicit Representation

CVPR 2021poster

We present a new pipeline for holistic 3D scene understanding from a single image, which could predict object shape, object pose and scene layout. As it is a highly ill-posed problem, existing methods usually suffer from inaccurate estimation of both shapes and layout especially for the cluttered sc…

Cited by 129PDFcodeScholar
2021

HumanGPS: Geodesic PreServing Feature for Dense Human Correspondences

CVPR 2021poster

In this paper, we address the problem of building pixel-wise dense correspondences between human images under arbitrary camera viewpoints and body poses. Previous methods either assume small motions or rely on discriminative descriptors extracted from local patches, which cannot handle large motion…

Cited by 14PDFScholar
2021

Interacting Two-Hand 3D Pose and Shape Reconstruction From Single Color Image

ICCV 2021poster

In this paper, we propose a novel deep learning framework to reconstruct 3D hand poses and shapes of two interacting hands from a single color image. Previous methods designed for single hand cannot be easily applied for the two hand scenario because of the heavy inter-hand occlusion and larger solu…

Cited by 111PDFcodeScholar
2021

Learning Compositional Representation for 4D Captures With Neural ODE

CVPR 2021poster

Learning based representation has become the key to the success of many computer vision systems. While many 3D representations have been proposed, it is still an unaddressed problem how to represent a dynamically changing 3D object. In this paper, we introduce a compositional representation for 4D c…

Cited by 33PDFScholar
2021

Learning Object-Compositional Neural Radiance Field for Editable Scene Rendering

ICCV 2021poster

Implicit neural rendering techniques have shown promising results for novel view synthesis. However, existing methods usually encode the entire scene as a whole, which is generally not aware of the object identity and limits the ability to the high-level editing tasks such as moving or adding furnit…

Cited by 356PDFScholar
2021

Multiresolution Deep Implicit Functions for 3D Shape Representation

ICCV 2021poster

We introduce Multiresolution Deep Implicit Functions (MDIF), a hierarchical representation that can recover fine geometry detail, while being able to perform global operations such as shape completion. Our model represents a complex 3D shape with a hierarchy of latent grids, which can be decoded int…

Cited by 52PDFScholar
2020

DIST: Rendering Deep Implicit Signed Distance Function With Differentiable Sphere Tracing

CVPR 2020poster

We propose a differentiable sphere tracing algorithm to bridge the gap between inverse graphics methods and the recently proposed deep learning based implicit signed distance function. Due to the nature of the implicit function, the rendering process requires tremendous function queries, which is pa…

Cited by 350PDFcodeScholar
2020

Deep Implicit Volume Compression

CVPR 2020oral

We describe a novel approach for compressing truncated signed distance fields (TSDF) stored in 3D voxel grids, and their corresponding textures. To compress the TSDF, our method relies on a block-based neural network architecture trained end-to-end, achieving state-of-the-art rate-distortion trade-o…

Cited by 52PDFcodeScholar
2020

DeepSFM: Structure From Motion Via Deep Bundle Adjustment

ECCV 2020poster

Structure from motion (SfM) is an essential computer vision problem which has not been well handled by deep learning. One of the promising trends is to apply explicit structural constraint, e.g. 3D cost volume, into the network. However, existing methods usually assume accurate camera poses either f…

Cited by 127SourcePDFScholar
2020

Du²Net: Learning Depth Estimation from Dual-Cameras and Dual-Pixels

ECCV 2020poster

Computational stereo has reached a high level of accuracy, but degrades in the presence of occlusions, repeated textures, and correspondence errors along edges. We present a novel approach based on neural networks for depth estimation that combines stereo from dual cameras with stereo from a dual-pi…

Cited by 39SourcePDFScholar
2020

GeoLayout: Geometry Driven Room Layout Estimation Based on Depth Maps of Planes

ECCV 2020poster

The task of room layout estimation is to locate the wall-floor, wall-ceiling, and wall-wall boundaries. Most recent methods solve this problem based on edge/keypoint detection or semantic segmentation. However, these approaches have shown limited attention on the geometry of the dominant planes and…

Cited by 34SourcePDFScholar
2020

Neural Pose Transfer by Spatially Adaptive Instance Normalization

CVPR 2020poster

Pose transfer has been studied for decades, in which the pose of a source mesh is applied to a target mesh. Particularly in this paper, we are interested in transferring the pose of source human mesh to deform the target human mesh, while the source and target meshes may have different identity info…

Cited by 71PDFcodeScholar
2019

DeepLiDAR: Deep Surface Normal Guided Depth Prediction for Outdoor Scene From Sparse LiDAR Data and Single Color Image

CVPR 2019poster

In this paper, we propose a deep learning architecture that produces accurate dense depth for the outdoor scene from a single color image and a sparse depth. Inspired by the indoor depth completion, our network estimates surface normals as the intermediate representation to produce dense depth, and…

Cited by 459PDFScholar
2018

ActiveStereoNet: End-to-End Self-Supervised Learning for Active Stereo Systems

ECCV 2018poster

In this paper we present ActiveStereoNet, the first deep learning solution for active stereo systems. Due to the lack of ground truth, our method is fully self-supervised, yet it produces precise depth with a subpixel precision of 1/30th of a pixel; it does not suffer from the common over-smoothing…

Cited by 139SourcePDFScholar
2018

Pixel2Mesh: Generating 3D Mesh Models from Single RGB Images

ECCV 2018poster

We propose an end-to-end deep learning architecture that produces a 3D shape in triangular mesh from a single color image. Limited by the nature of deep neural network, previous methods usually represent a 3D shape in volume or point cloud, and it is non-trivial to convert them to the more ready-to-…

Cited by 1707SourcePDFScholar
2017

DeepContext: Context-Encoding Neural Pathways for 3D Holistic Scene Understanding

ICCV 2017poster

3D context has been shown to be an extremely important cue for scene understanding, yet very little research has been done on integrating context information with deep models. This paper presents an approach to embed 3D context into the topology of a neural network trained to perform holistic scene…

Cited by 82PDFScholar
2017

Physically-Based Rendering for Indoor Scene Understanding Using Convolutional Neural Networks

CVPR 2017poster

Indoor scene understanding is central to applications such as robot navigation and human companion assistance. Over the last years, data-driven deep neural networks have outperformed many traditional approaches thanks to their representation learning capabilities. One of the bottlenecks in training…

Cited by 329PDFScholar