← Search

Shunsuke Saito

44 accepted papers

2026

GeoRelight: Learning Joint Geometrical Relighting and Reconstruction with Flexible Multi-Modal Diffusion Transformers

CVPR 2026

Relighting a person from a single photo is an attractive but ill-posed task, as a 2D image ambiguously entangles 3D geometry, intrinsic appearance, and illumination. Current methods either use sequential pipelines that suffer from error accumulation, or they do not explicitly leverage 3D geometry du

Cited by 0SourceScholar
2026

Large-scale Codec Avatars: The Unreasonable Effectiveness of Large-scale Avatar Pretraining

CVPR 2026

High-quality 3D avatar modeling faces a critical trade-off between fidelity and generalization. On the one hand, multi-view studio data enables high-fidelity modeling of humans with precise control over expressions and poses, but it struggles to generalize to real-world data due to limited scale and

Cited by 0SourcecodeScholar
2025

3DTopia-XL: Scaling High-quality 3D Asset Generation via Primitive Diffusion

CVPR 2025highlight

The increasing demand for high-quality 3D assets across various industries necessitates efficient and automated 3D content creation. Despite recent advancements in 3D generative models, existing methods still face challenges with optimization speed, geometric fidelity, and the lack of assets for phy…

2025

ATLAS: Decoupling Skeletal and Shape Parameters for Expressive Parametric Human Modeling

ICCV 2025poster

Parametric body models offer expressive 3D representation of humans across a wide range of poses, shapes, and facial expressions, typically derived by learning a basis over registered 3D meshes. However, existing human mesh modeling approaches struggle to capture detailed variations across diverse b…

Cited by 0SourcePDFScholar
2025

Agent-to-Sim: Learning Interactive Behavior Models from Casual Longitudinal Videos

ICLR 2025poster

We present Agent-to-Sim (ATS), a framework for learning interactive behavior models of 3D agents from casual longitudinal video collections. Different from prior works that rely on marker-based tracking and multiview cameras, ATS learns natural behaviors of animal agents non-invasively through video…

2025

Avat3r: Large Animatable Gaussian Reconstruction Model for High-fidelity 3D Head Avatars

ICCV 2025poster

Traditionally, creating photo-realistic 3D head avatars requires a studio-level multi-view capture setup and expensive optimization during test-time, limiting the use of digital human doubles to the VFX industry or offline renderings. To address this shortcoming, we present Avat3r, which regresses a…

Cited by 0SourcePDFScholar
2025

FRESA: Feedforward Reconstruction of Personalized Skinned Avatars from Few Images

CVPR 2025highlight

We present a novel method for reconstructing personalized 3D human avatars with realistic animation from only a few images. Due to the large variations in body shapes, poses, and cloth types, existing methods mostly require hours of per-subject optimization during inference, which limits their pract…

2025

Generative Modeling of Shape-Dependent Self-Contact Human Poses

ICCV 2025poster

One can hardly model self-contact of human poses without considering underlying body shapes. For example, the pose of rubbing a belly for a person with a low BMI leads to penetration of the hand into the belly for a person with a high BMI. Despite its relevance, existing self-contact datasets lack t…

Cited by 0SourcePDFScholar
2025

HairCUP: Hair Compositional Universal Prior for 3D Gaussian Avatars

ICCV 2025poster

We present a universal prior model for 3D head avatars with explicit hair compositionality. Existing approaches to build generalizable priors for 3D head avatars often adopt a holistic modeling approach, treating the face and hair as an inseparable entity. This overlooks the inherent compositionalit…

Cited by 0SourcePDFScholar
2025

PGC: Physics-Based Gaussian Cloth from a Single Pose

CVPR 2025highlight

We introduce a novel approach to reconstruct simulation-ready garments with intricate appearance. Despite recent advancements, existing methods often struggle to balance the need for accurate garment reconstruction with the ability to generalize to new poses and body shapes or require large amounts…

2025

Pippo: High-Resolution Multi-View Humans from a Single Image

CVPR 2025highlight

We present Pippo, a generative model capable of producing 1K resolution dense turnaround videos of a person from a single casually clicked photo. Pippo is a multi-view diffusion transformer and does not require any additional inputs - e.g., a fitted parametric model or camera parameters of the input…

Cited by 1SourcePDFScholar
2025

REWIND: Real-Time Egocentric Whole-Body Motion Diffusion with Exemplar-Based Identity Conditioning

CVPR 2025poster

We present REWIND (Real-Time Egocentric Whole-Body Motion Diffusion), a one-step diffusion model for real-time, high-fidelity human motion estimation from egocentric image inputs. While an existing method for egocentric whole-body (i.e., body and hands) motion estimation is non-real-time and acausal…

Cited by 0SourcePDFScholar
2025

Vid2Avatar-Pro: Authentic Avatar from Videos in the Wild via Universal Prior

CVPR 2025poster

We present Vid2Avatar-Pro, a method to create photorealistic and animatable 3D human avatars from monocular in-the-wild videos. Building a high-quality avatar that supports animation with diverse poses from a monocular video is challenging because the observation of pose diversity and view points is…

Cited by 0SourcePDFScholar
2024

Bridging the Gap: Studio-like Avatar Creation from a Monocular Phone Capture

ECCV 2024oral

"Creating photorealistic avatars for individuals traditionally involves extensive capture sessions with complex and expensive studio devices like the LightStage system. While recent strides in neural representations have enabled the generation of photorealistic and animatable 3D avatars from quick p…

Cited by 0SourcePDFScholar
2024

Codec Avatar Studio: Paired Human Captures for Complete, Driveable, and Generalizable Avatars

NeurIPS 2024poster

To build photorealistic avatars that users can embody, human modelling must be complete (cover the full body), driveable (able to reproduce the current motion and appearance from the user), and generalizable (_i.e._, easily adaptable to novel identities). Towards these goals, _paired_ captures, that…

2024

GALA: Generating Animatable Layered Assets from a Single Scan

CVPR 2024poster

We present GALA a framework that takes as input a single-layer clothed 3D human mesh and decomposes it into complete multi-layered 3D assets. The outputs can then be combined with other assets to create novel clothed human avatars with any pose. Existing reconstruction approaches often treat clothed…

Cited by 7SourcePDFScholar
2024

High-Fidelity Modeling of Generalizable Wrinkle Deformation

ECCV 2024poster

"This paper proposes a generalizable model to synthesize high-fidelity clothing wrinkle deformation in 3D by learning from real data. Given the complex deformation behaviors of real-world clothing, this task presents significant challenges, primarily due to the lack of accurate ground-truth data. Ob…

Cited by 0SourcePDFScholar
2024

InterHandGen: Two-Hand Interaction Generation via Cascaded Reverse Diffusion

CVPR 2024poster

We present InterHandGen a novel framework that learns the generative prior of two-hand interaction. Sampling from our model yields plausible and diverse two-hand shapes in close interaction with or without an object. Our prior can be incorporated into any optimization or learning methods to reduce a…

Cited by 13SourcePDFScholar
2024

Sapiens: Foundation for Human Vision Models

ECCV 2024oral

"We present Sapiens, a family of models for four fundamental human-centric vision tasks – 2D pose estimation, body-part segmentation, depth estimation, and surface normal prediction. Our models natively support 1K high-resolution inference and are extremely easy to adapt for individual tasks by simp…

Cited by 28SourcePDFScholar
2024

URHand: Universal Relightable Hands

CVPR 2024poster

Existing photorealistic relightable hand models require extensive identity-specific observations in different views poses and illuminations and face challenges in generalizing to natural illuminations and novel identities. To bridge this gap we present URHand the first universal relightable hand mod…

Cited by 11SourcePDFScholar
2023

A Dataset of Relighted 3D Interacting Hands

NeurIPS 2023poster

The two-hand interaction is one of the most challenging signals to analyze due to the self-similarity, complicated articulations, and occlusions of hands. Although several datasets have been proposed for the two-hand interaction analysis, all of them do not achieve 1) diverse and realistic image app…

2023

MEGANE: Morphable Eyeglass and Avatar Network

CVPR 2023poster

Eyeglasses play an important role in the perception of identity. Authentic virtual representations of faces can benefit greatly from their inclusion. However, modeling the geometric and appearance interactions of glasses and the face of virtual representations of humans is challenging. Glasses and f…

Cited by 16SourcePDFScholar
2023

NCHO: Unsupervised Learning for Neural 3D Composition of Humans and Objects

ICCV 2023poster

Deep generative models have been recently extended to synthesizing 3D digital humans. However, previous approaches treat clothed humans as a single chunk of geometry without considering the compositionality of clothing and accessories. As a result, individual items cannot be naturally composed into…

Cited by 12PDFcodeScholar
2023

RelightableHands: Efficient Neural Relighting of Articulated Hand Models

CVPR 2023poster

We present the first neural relighting approach for rendering high-fidelity personalized hands that can be animated in real-time under novel illumination. Our approach adopts a teacher-student framework, where the teacher learns appearance under a single point light from images captured in a light-s…

Cited by 17SourcePDFScholar
2022

AutoAvatar: Autoregressive Neural Fields for Dynamic Avatar Modeling

ECCV 2022poster

"Neural fields such as implicit surfaces have recently enabled avatar modeling from raw scans without explicit temporal correspondences. In this work, we exploit autoregressive modeling to further extend this notion to capture dynamic effects, such as soft-tissue deformations. Although autoregressiv…

2022

COAP: Compositional Articulated Occupancy of People

CVPR 2022poster

We present a novel neural implicit representation for articulated human bodies. Compared to explicit template meshes, neural implicit body representations provide an efficient mechanism for modeling interactions with the environment, which is essential for human motion reconstruction and synthesis i…

Cited by 57PDFcodeScholar
2022

KeypointNeRF: Generalizing Image-Based Volumetric Avatars Using Relative Spatial Encoding of Keypoints

ECCV 2022poster

"Image-based volumetric avatars using pixel-aligned features promise generalization to unseen poses and identities. Prior work leverages global spatial encodings and multi-view geometric consistency to reduce spatial ambiguity. However, global encodings often suffer from overfitting to the distribut…

2022

Neural Strands: Learning Hair Geometry and Appearance from Multi-View Images

ECCV 2022poster

"We present Neural Strands, a novel learning framework for modeling accurate hair geometry and appearance from multi-view image inputs. The learned hair model can be rendered in real-time from any viewpoint with high-fidelity view-dependent effects. Our model achieves intuitive shape and style contr…

Cited by 44SourcePDFScholar
2021

ARCH++: Animation-Ready Clothed Human Reconstruction Revisited

ICCV 2021poster

We present ARCH++, an image-based method to reconstruct 3D avatars with arbitrary clothing styles. Our reconstructed avatars are animation-ready and highly realistic, in both the visible regions from input views and the unseen regions. While prior work shows great promise of reconstructing animatabl…

Cited by 222PDFScholar
2021

SCALE: Modeling Clothed Humans with a Surface Codec of Articulated Local Elements

CVPR 2021poster

Learning to model and reconstruct humans in clothing is challenging due to articulation, non-rigid deformation, and varying clothing types and topologies. To enable learning, the choice of representation is the key. Recent work uses neural networks to parameterize local surface elements. This approa…

Cited by 114PDFcodeScholar
2021

SCANimate: Weakly Supervised Learning of Skinned Clothed Avatar Networks

CVPR 2021poster

We present SCANimate, an end-to-end trainable framework that takes raw 3D scans of a clothed human and turns them into an animatable avatar. These avatars are driven by pose parameters and have realistic clothing that moves and deforms naturally. SCANimate does not rely on a customized mesh template…

Cited by 264PDFcodeScholar
2020

Monocular Real-Time Volumetric Performance Capture

ECCV 2020poster

We present the first approach to volumetric performance capture and novel-view rendering at real-time speed from monocular video, eliminating the need for expensive multi-view systems or cumbersome pre-acquisition of a personalized template model. Our system reconstructs a fully textured 3D human fr…

2020

PIFuHD: Multi-Level Pixel-Aligned Implicit Function for High-Resolution 3D Human Digitization

CVPR 2020oral

Recent advances in image-based 3D human shape estimation have been driven by the significant improvement in representation power afforded by deep neural networks. Although current approaches have demonstrated the potential in real world settings, they still fail to produce reconstructions with the l…

Cited by 955PDFcodeScholar
2019

PIFu: Pixel-Aligned Implicit Function for High-Resolution Clothed Human Digitization

ICCV 2019poster

We introduce Pixel-aligned Implicit Function (PIFu), an implicit representation that locally aligns pixels of 2D images with the global context of their corresponding 3D object. Using PIFu, we propose an end-to-end deep learning method for digitizing highly detailed clothed humans that can infer bot…

Cited by 1438PDFScholar
2018

Mesoscopic Facial Geometry Inference Using Deep Neural Networks

CVPR 2018poster

We present a learning-based approach for synthesizing facial geometry at medium and fine scales from diffusely-lit facial texture maps. When applied to an image sequence, the synthesized detail is temporally coherent. Unlike current state-of-the-art methods, which assume "dark is deep", our model…

Cited by 76SourcePDFScholar
2017

Learning Dense Facial Correspondences in Unconstrained Images

ICCV 2017poster

We present a minimalistic but effective neural network that computes dense facial correspondences in highly unconstrained RGB images. Our network learns a per-pixel flow and a matchability mask between 2D input photographs of a person and the projection of a textured 3D face model. To train such a n…

Cited by 80PDFScholar
2017

Photorealistic Facial Texture Inference Using Deep Neural Networks

CVPR 2017spotlight

We present a data-driven inference method that can synthesize a photorealistic texture map of a complete 3D face model given a partial 2D view of a person in the wild. After an initial estimation of shape and low-frequency albedo, we compute a high-frequency partial texture map, without the shading…

Cited by 162PDFcodeScholar
2017

Realistic Dynamic Facial Textures From a Single Image Using GANs

ICCV 2017poster

We present a novel method to realistically puppeteer and animate a face from a single RGB image using a source video sequence. We begin by fitting a multilinear PCA model to obtain the 3D geometry and a single texture of the target face. In order for the animation to be realistic, however, we need d…

Cited by 116PDFScholar