← Search

Tomas Simon

20 accepted papers

2025

HairCUP: Hair Compositional Universal Prior for 3D Gaussian Avatars

ICCV 2025poster

We present a universal prior model for 3D head avatars with explicit hair compositionality. Existing approaches to build generalizable priors for 3D head avatars often adopt a holistic modeling approach, treating the face and hair as an inseparable entity. This overlooks the inherent compositionalit…

Cited by 0SourcePDFScholar
2024

Codec Avatar Studio: Paired Human Captures for Complete, Driveable, and Generalizable Avatars

NeurIPS 2024poster

To build photorealistic avatars that users can embody, human modelling must be complete (cover the full body), driveable (able to reproduce the current motion and appearance from the user), and generalizable (_i.e._, easily adaptable to novel identities). Towards these goals, _paired_ captures, that…

2024

Rasterized Edge Gradients: Handling Discontinuities Differentially

ECCV 2024oral

"Computing the gradients of a rendering process is paramount for diverse applications in computer vision and graphics. However, accurate computation of these gradients is challenging due to discontinuities and rendering approximations, particularly for surface-based representations and rasterization…

Cited by 3SourcePDFScholar
2024

URHand: Universal Relightable Hands

CVPR 2024poster

Existing photorealistic relightable hand models require extensive identity-specific observations in different views poses and illuminations and face challenges in generalizing to natural illuminations and novel identities. To bridge this gap we present URHand the first universal relightable hand mod…

Cited by 11SourcePDFScholar
2023

A Dataset of Relighted 3D Interacting Hands

NeurIPS 2023poster

The two-hand interaction is one of the most challenging signals to analyze due to the self-similarity, complicated articulations, and occlusions of hands. Although several datasets have been proposed for the two-hand interaction analysis, all of them do not achieve 1) diverse and realistic image app…

2023

MEGANE: Morphable Eyeglass and Avatar Network

CVPR 2023poster

Eyeglasses play an important role in the perception of identity. Authentic virtual representations of faces can benefit greatly from their inclusion. However, modeling the geometric and appearance interactions of glasses and the face of virtual representations of humans is challenging. Glasses and f…

Cited by 16SourcePDFScholar
2023

RelightableHands: Efficient Neural Relighting of Articulated Hand Models

CVPR 2023poster

We present the first neural relighting approach for rendering high-fidelity personalized hands that can be animated in real-time under novel illumination. Our approach adopts a teacher-student framework, where the teacher learns appearance under a single point light from images captured in a light-s…

Cited by 17SourcePDFScholar
2021

Learning Compositional Radiance Fields of Dynamic Human Heads

CVPR 2021poster

Photorealistic rendering of dynamic humans is an important ability for telepresence systems, virtual shopping, synthetic data generation, and more. Recently, neural rendering methods, which combine techniques from computer graphics and machine learning, have created high-fidelity models of humans an…

Cited by 101PDFScholar
2021

SimPoE: Simulated Character Control for 3D Human Pose Estimation

CVPR 2021poster

Accurate estimation of 3D human motion from monocular video requires modeling both kinematics (body motion without physical forces) and dynamics (motion with physical forces). To demonstrate this, we present SimPoE, a Simulation-based approach for 3D human Pose Estimation, which integrates image-bas…

Cited by 163PDFScholar
2020

PIFuHD: Multi-Level Pixel-Aligned Implicit Function for High-Resolution 3D Human Digitization

CVPR 2020oral

Recent advances in image-based 3D human shape estimation have been driven by the significant improvement in representation power afforded by deep neural networks. Although current approaches have demonstrated the potential in real world settings, they still fail to produce reconstructions with the l…

Cited by 955PDFcodeScholar
2019

LBS Autoencoder: Self-Supervised Fitting of Articulated Meshes to Point Clouds

CVPR 2019poster

We present LBS-AE; a self-supervised autoencoding algorithm for fitting articulated mesh models to point clouds. As input, we take a sequence of point clouds to be registered as well as an artist-rigged mesh, i.e. a template mesh equipped with a linear-blend skinning (LBS) deformation space paramete…

Cited by 51PDFScholar
2019

Single-Network Whole-Body Pose Estimation

ICCV 2019poster

We present the first single-network approach for 2D whole-body pose estimation, which entails simultaneous localization of body, face, hands, and feet keypoints. Due to the bottom-up formulation, our method maintains constant real-time performance regardless of the number of people in the image. The…

Cited by 135PDFcodeScholar
2019

Towards Social Artificial Intelligence: Nonverbal Social Signal Prediction in a Triadic Interaction

CVPR 2019oral

We present a new research task and a dataset to understand human social interactions via computational methods, to ultimately endow machines with the ability to encode and decode a broad channel of social signals humans use. This research direction is essential to make a machine that genuinely commu…

Cited by 116PDFcodeScholar
2018

Total Capture: A 3D Deformation Model for Tracking Faces, Hands, and Bodies

CVPR 2018poster

We present a unified deformation model for the markerless capture of multiple scales of human movement, including facial expressions, body motion, and hand gestures. An initial model is generated by locally stitching together models of the individual parts of the human body, which we refer to as the…

Cited by 617SourcePDFScholar
2017

Hand Keypoint Detection in Single Images Using Multiview Bootstrapping

CVPR 2017poster

We present an approach that uses a multi-camera system to train fine-grained detectors for keypoints that are prone to occlusion, such as the joints of a hand. We call this procedure multiview bootstrapping: first, an initial keypoint detector is used to produce noisy labels in multiple views of the…

Cited by 1579PDFScholar
2017

Realtime Multi-Person 2D Pose Estimation Using Part Affinity Fields

CVPR 2017oral

We present an approach to efficiently detect the 2D pose of multiple people in an image. The approach uses a nonparametric representation, which we refer to as Part Affinity Fields (PAFs), to learn to associate body parts with individuals in the image. The architecture encodes global context, allowi…

Cited by 9379PDFcodeScholar
2015

Photogeometric Scene Flow for High-Detail Dynamic 3D Reconstruction

ICCV 2015poster

Photometric stereo (PS) is an established technique for high-detail reconstruction of 3D geometry and appearance. To correct for surface integration errors, PS is often combined with multiview stereo (MVS). With dynamic objects, PS reconstruction also faces the problem of computing optical flow (OF)…

Cited by 61PDFScholar