← Search

Yaser Sheikh

38 accepted papers

2026

Large-scale Codec Avatars: The Unreasonable Effectiveness of Large-scale Avatar Pretraining

CVPR 2026

High-quality 3D avatar modeling faces a critical trade-off between fidelity and generalization. On the one hand, multi-view studio data enables high-fidelity modeling of humans with precise control over expressions and poses, but it struggles to generalize to real-world data due to limited scale and

Cited by 0SourcecodeScholar
2025

FRESA: Feedforward Reconstruction of Personalized Skinned Avatars from Few Images

CVPR 2025highlight

We present a novel method for reconstructing personalized 3D human avatars with realistic animation from only a few images. Due to the large variations in body shapes, poses, and cloth types, existing methods mostly require hours of per-subject optimization during inference, which limits their pract…

2025

Vid2Avatar-Pro: Authentic Avatar from Videos in the Wild via Universal Prior

CVPR 2025poster

We present Vid2Avatar-Pro, a method to create photorealistic and animatable 3D human avatars from monocular in-the-wild videos. Building a high-quality avatar that supports animation with diverse poses from a monocular video is challenging because the observation of pose diversity and view points is…

Cited by 0SourcePDFScholar
2024

Codec Avatar Studio: Paired Human Captures for Complete, Driveable, and Generalizable Avatars

NeurIPS 2024poster

To build photorealistic avatars that users can embody, human modelling must be complete (cover the full body), driveable (able to reproduce the current motion and appearance from the user), and generalizable (_i.e._, easily adaptable to novel identities). Towards these goals, _paired_ captures, that…

2024

Rasterized Edge Gradients: Handling Discontinuities Differentially

ECCV 2024oral

"Computing the gradients of a rendering process is paramount for diverse applications in computer vision and graphics. However, accurate computation of these gradients is challenging due to discontinuities and rendering approximations, particularly for surface-based representations and rasterization…

Cited by 3SourcePDFScholar
2024

URHand: Universal Relightable Hands

CVPR 2024poster

Existing photorealistic relightable hand models require extensive identity-specific observations in different views poses and illuminations and face challenges in generalizing to natural illuminations and novel identities. To bridge this gap we present URHand the first universal relightable hand mod…

Cited by 11SourcePDFScholar
2023

RelightableHands: Efficient Neural Relighting of Articulated Hand Models

CVPR 2023poster

We present the first neural relighting approach for rendering high-fidelity personalized hands that can be animated in real-time under novel illumination. Our approach adopts a teacher-student framework, where the teacher learns appearance under a single point light from images captured in a light-s…

Cited by 17SourcePDFScholar
2021

High-Fidelity Face Tracking for AR/VR via Deep Lighting Adaptation

CVPR 2021poster

3D video avatars can empower virtual communications by providing compression, privacy, entertainment, and a sense of presence in AR/VR. Best 3D photo-realistic AR/VR avatars driven by video, that can minimize uncanny effects, rely on person-specific models. However, existing person-specific photo-re…

Cited by 29PDFScholar
2021

Implicit HRTF Modeling Using Temporal Convolutional Networks

ICASSP 2021accepted

Estimation of accurate head-related transfer functions (HRTFs) is crucial to achieve realistic binaural acoustic experiences. HRTFs depend on source/listener locations and are therefore expensive and cumbersome to measure; traditional approaches require listener-dependent measurements of HRTFs at th…

Cited by 0SourceScholar
2021

MeshTalk: 3D Face Animation From Speech Using Cross-Modality Disentanglement

ICCV 2021poster

This paper presents a generic method for generating full facial 3D animation from speech. Existing approaches to audio-driven facial animation exhibit uncanny or static upper face animation, fail to produce accurate and plausible co-articulation or rely on person-specific models that limit their sca…

Cited by 241PDFcodeScholar
2021

Neural Synthesis of Binaural Speech From Mono Audio

ICLR 2021oral

We present a neural rendering approach for binaural sound synthesis that can produce realistic and spatially accurate binaural sound in realtime. The network takes, as input, a single-channel audio source and synthesizes, as output, two-channel binaural sound, conditioned on the relative position an…

Cited by 76SourcePDFScholar
2020

4D Visualization of Dynamic Events From Unconstrained Multi-View Videos

CVPR 2020poster

We present a data-driven approach for 4D space-time visualization of dynamic events from videos captured by hand-held multiple cameras. Key to our approach is the use of self-supervised neural networks specific to the scene to compose static and dynamic aspects of an event. Though captured from disc…

Cited by 83PDFScholar
2020

Expressive Telepresence via Modular Codec Avatars

ECCV 2020poster

VR telepresence consists of interacting with another human in a virtual space represented by an avatar. Today most avatars are cartoon-like, but soon the technology will allow video-realistic ones. This paper aims in this direction and presents Modular Codec Avatars (MCA), a method to generate hyper…

Cited by 40SourcePDFScholar
2020

Fully Convolutional Mesh Autoencoder using Efficient Spatially Varying Kernels

NeurIPS 2020poster

Learning latent representations of registered meshes is useful for many 3D tasks. Techniques have recently shifted to neural mesh autoencoders. Although they demonstrate higher precision than traditional methods, they remain unable to capture fine-grained deformations. Furthermore, these methods can…

Cited by 97SourcePDFScholar
2019

Efficient Online Multi-Person 2D Pose Tracking With Recurrent Spatio-Temporal Affinity Fields

CVPR 2019oral

We present an online approach to efficiently and simultaneously detect and track 2D poses of multiple people in a video sequence. We build upon Part Affinity Field (PAF) representation designed for static images, and propose an architecture that can encode and predict Spatio-Temporal Affinity Fields…

Cited by 153PDFScholar
2019

LBS Autoencoder: Self-Supervised Fitting of Articulated Meshes to Point Clouds

CVPR 2019poster

We present LBS-AE; a self-supervised autoencoding algorithm for fitting articulated mesh models to point clouds. As input, we take a sequence of point clouds to be registered as well as an artist-rigged mesh, i.e. a template mesh equipped with a linear-blend skinning (LBS) deformation space paramete…

Cited by 51PDFScholar
2019

Single-Network Whole-Body Pose Estimation

ICCV 2019poster

We present the first single-network approach for 2D whole-body pose estimation, which entails simultaneous localization of body, face, hands, and feet keypoints. Due to the bottom-up formulation, our method maintains constant real-time performance regardless of the number of people in the image. The…

Cited by 135PDFcodeScholar
2019

Talking With Hands 16.2M: A Large-Scale Dataset of Synchronized Body-Finger Motion and Audio for Conversational Motion Analysis and Synthesis

ICCV 2019poster

We present a 16.2-million frame (50-hour) multimodal dataset of two-person face-to-face spontaneous conversations. Our dataset features synchronized body and finger motion as well as audio data. To the best of our knowledge, it represents the largest motion capture and audio dataset of natural conve…

Cited by 118PDFScholar
2019

Towards Social Artificial Intelligence: Nonverbal Social Signal Prediction in a Triadic Interaction

CVPR 2019oral

We present a new research task and a dataset to understand human social interactions via computational methods, to ultimately endow machines with the ability to encode and decode a broad channel of social signals humans use. This research direction is essential to make a machine that genuinely commu…

Cited by 116PDFcodeScholar
2018

Learning Patch Reconstructability for Accelerating Multi-View Stereo

CVPR 2018poster

We present an approach to accelerate multi-view stereo (MVS) by prioritizing computation on image patches that are likely to produce accurate 3D surface reconstructions. Our key insight is that the accuracy of the surface reconstruction from a given image patch can be predicted significantly faster…

Cited by 9SourcePDFScholar
2018

Modeling Facial Geometry Using Compositional VAEs

CVPR 2018poster

We propose a method for learning non-linear face geometry representations using deep generative models. Our model is a variational autoencoder with multiple levels of hidden variables where lower layers capture global geometry and higher ones encode more local deformations. Based on that, we pr…

Cited by 151SourcePDFScholar
2018

Supervision-by-Registration: An Unsupervised Approach to Improve the Precision of Facial Landmark Detectors

CVPR 2018poster

In this paper, we present supervision-by-registration, an unsupervised approach to improve the precision of facial landmark detectors on both images and video. Our key observation is that the detections of the same landmark in adjacent frames should be coherent with registration, i.e., optical flow.…

Cited by 252SourcePDFScholar
2018

Total Capture: A 3D Deformation Model for Tracking Faces, Hands, and Bodies

CVPR 2018poster

We present a unified deformation model for the markerless capture of multiple scales of human movement, including facial expressions, body motion, and hand gestures. An initial model is generated by locally stitching together models of the individual parts of the human body, which we refer to as the…

Cited by 617SourcePDFScholar
2017

Hand Keypoint Detection in Single Images Using Multiview Bootstrapping

CVPR 2017poster

We present an approach that uses a multi-camera system to train fine-grained detectors for keypoints that are prone to occlusion, such as the joints of a hand. We call this procedure multiview bootstrapping: first, an initial keypoint detector is used to produce noisy labels in multiple views of the…

Cited by 1579PDFScholar
2017

Realtime Multi-Person 2D Pose Estimation Using Part Affinity Fields

CVPR 2017oral

We present an approach to efficiently detect the 2D pose of multiple people in an image. The approach uses a nonparametric representation, which we refer to as Part Affinity Fields (PAFs), to learn to associate body parts with individuals in the image. The architecture encodes global context, allowi…

Cited by 9379PDFcodeScholar
2015

Panoptic Studio: A Massively Multiview System for Social Motion Capture

ICCV 2015oral

We present an approach to capture the 3D structure and motion of a group of people engaged in a social interaction. The core challenges in capturing social interactions are: (1) occlusion is functional and frequent; (2) subtle motion needs to be measured over a space large enough to host a social gr…

Cited by 1063PDFScholar
2015

Photogeometric Scene Flow for High-Detail Dynamic 3D Reconstruction

ICCV 2015poster

Photometric stereo (PS) is an established technique for high-detail reconstruction of 3D geometry and appearance. To correct for surface integration errors, PS is often combined with multiview stereo (MVS). With dynamic objects, PS reconstruction also faces the problem of computing optical flow (OF)…

Cited by 61PDFScholar