← Search

Mihai Zanfir

14 accepted papers

2023

Structured 3D Features for Reconstructing Controllable Avatars

CVPR 2023poster

We introduce Structured 3D Features, a model based on a novel implicit 3D representation that pools pixel-aligned image features onto dense 3D points sampled from a parametric, statistical human mesh surface. The 3D points have associated semantics and can move freely in 3D space. This allows for op…

Cited by 39SourcePDFScholar
2022

HUM3DIL: Semi-supervised Multi-modal 3D HumanPose Estimation for Autonomous Driving

CoRL 2022poster

Autonomous driving is an exciting new industry, posing important research questions. Within the perception module, 3D human pose estimation is an emerging technology, which can enable the autonomous vehicle to perceive and understand the subtle and complex behaviors of pedestrians. While hardware sy…

Cited by 32SourceScholar
2022

Photorealistic Monocular 3D Reconstruction of Humans Wearing Clothing

CVPR 2022poster

We present PHORHUM, a novel, end-to-end trainable, deep neural network methodology for photorealistic 3D human reconstruction given just a monocular RGB image. Our pixel-aligned method estimates detailed 3D geometry and, for the first time, the unshaded surface color together with the scene illumina…

Cited by 161PDFScholar
2021

AIFit: Automatic 3D Human-Interpretable Feedback Models for Fitness Training

CVPR 2021poster

I went to the gym today, but how well did I do? And where should I improve? Ah, my back hurts slightly... User engagement can be sustained and injuries avoided by being able to reconstruct 3d human pose and motion, relate it to good training practices, identify errors, and provide early, real-time f…

Cited by 92PDFScholar
2021

Learning Complex 3D Human Self-Contact

AAAI 2021technical

Monocular estimation of three dimensional human self-contact is fundamental for detailed scene analysis including body language understanding and behaviour modeling. Existing 3d reconstruction methods do not focus on body regions in self-contact and consequently recover configurations that are eith…

Cited by 33SourcePDFScholar
2021

Neural Descent for Visual 3D Human Pose and Shape

CVPR 2021poster

We present deep neural network methodology to reconstruct the 3d pose and shape of people, including hand gestures and facial expression, given an input RGB image. We rely on a recently introduced, expressive full body statistical 3d human model, GHUM, trained end-to-end, and learn to reconstruct it…

Cited by 75PDFScholar
2021

REMIPS: Physically Consistent 3D Reconstruction of Multiple Interacting People under Weak Supervision

NeurIPS 2021poster

The three-dimensional reconstruction of multiple interacting humans given a monocular image is crucial for the general task of scene understanding, as capturing the subtleties of interaction is often the very reason for taking a picture. Current 3D human reconstruction methods either treat each pers…

Cited by 30SourcePDFScholar
2021

THUNDR: Transformer-Based 3D Human Reconstruction With Markers

ICCV 2021poster

We present THUNDR, a transformer-based deep neural network methodology to reconstruct the 3d pose and shape of people, given monocular RGB images. Key to our methodology is an intermediate 3d marker representation, where we aim to combine the predictive power of model-free-output architectures and t…

Cited by 83PDFScholar
2020

Three-Dimensional Reconstruction of Human Interactions

CVPR 2020poster

Understanding 3d human interactions is fundamental for fine grained scene analysis and behavioural modeling. However, most of the existing models focus on analyzing a single person in isolation, and those who process several people focus largely on resolving multi-person data association, rather tha…

Cited by 131PDFScholar
2018

3D Human Sensing, Action and Emotion Recognition in Robot Assisted Therapy of Children With Autism

CVPR 2018poster

We introduce new, fine-grained action and emotion recognition tasks defined on non-staged videos, recorded during robot-assisted therapy sessions of children with autism. The tasks present several challenges: a large dataset with long videos, a large number of highly variable actions, children that…

Cited by 137SourcePDFScholar
2018

Deep Network for the Integrated 3D Sensing of Multiple People in Natural Images

NeurIPS 2018spotlight

We present MubyNet -- a feed-forward, multitask, bottom up system for the integrated localization, as well as 3d pose and shape estimation, of multiple people in monocular images. The challenge is the formal modeling of the problem that intrinsically requires discrete and continuous computation, e.g…

Cited by 170SourcePDFScholar