← Search

Cristian Sminchisescu

51 accepted papers

2025

VLOGGER: Multimodal Diffusion for Embodied Avatar Synthesis

CVPR 2025poster

We propose VLOGGER, a method for audio-driven human video generation from a single input image of a person, which builds on the success of recent generative diffusion models. Our method consists of 1) a stochastic human-to-3d-motion diffusion model, and 2) a novel diffusion-based architecture that a…

Cited by 27SourcePDFScholar
2024

DiffHuman: Probabilistic Photorealistic 3D Reconstruction of Humans

CVPR 2024poster

We present DiffHuman a probabilistic method for photorealistic 3D human reconstruction from a single RGB image. Despite the ill-posed nature of this problem most methods are deterministic and output a single solution often resulting in a lack of geometric detail and blurriness in unseen or uncertain…

Cited by 5SourcePDFScholar
2024

Instant 3D Human Avatar Generation using Image Diffusion Models

ECCV 2024poster

"We present , a method for fast, high quality 3D human avatar generation from different input modalities, such as images and text prompts and with control over the generated pose and shape. The common theme is the use of diffusion-based image generation networks that are specialized for each particu…

Cited by 5SourcePDFScholar
2024

Learned Neural Physics Simulation for Articulated 3D Human Pose Reconstruction

ECCV 2024poster

"We propose a novel neural network approach to model the dynamics of articulated human motion with contact. Our goal is to develop a faster and more convenient alternative to traditional physics simulators for use in computer vision tasks such as human motion reconstruction from video. To that end w…

Cited by 2SourcePDFScholar
2024

Score Distillation Sampling with Learned Manifold Corrective

ECCV 2024poster

"Score Distillation Sampling (SDS) is a recent but already widely popular method that relies on an image diffusion model to control optimization problems using text prompts. aIn this paper, we conduct an in-depth analysis of the SDS loss function, identify an inherent problem with its formulation, a…

Cited by 6SourcePDFScholar
2023

DreamHuman: Animatable 3D Avatars from Text

NeurIPS 2023spotlight

We present \emph{DreamHuman}, a method to generate realistic animatable 3D human avatar models entirely from textual descriptions. Recent text-to-3D methods have made considerable strides in generation, but are still lacking in important aspects. Control and often spatial resolution remain limited,…

Cited by 96SourcePDFScholar
2023

Structured 3D Features for Reconstructing Controllable Avatars

CVPR 2023poster

We introduce Structured 3D Features, a model based on a novel implicit 3D representation that pools pixel-aligned image features onto dense 3D points sampled from a parametric, statistical human mesh surface. The 3D points have associated semantics and can move freely in 3D space. This allows for op…

Cited by 39SourcePDFScholar
2023

Transformer-Based Learned Optimization

CVPR 2023poster

We propose a new approach to learned optimization where we represent the computation of an optimizer's update step using a neural network. The parameters of the optimizer are then learned by training on a set of optimization tasks with the objective to perform minimization efficiently. Our innovatio…

2022

A Real-Time Online Learning Framework for Joint 3D Reconstruction and Semantic Segmentation of Indoor Scenes

RA-L 2022

This letter presents a real-time online vision framework to jointly recover an indoor scene’s 3D structure and semantic label. Given noisy depth maps, a camera trajectory, and 2D semantic labels at train time, the proposed deep neural network based approach learns to fuse the depth over frames with

Cited by 26SourcecodeScholar
2022

BEHAVE: Dataset and Method for Tracking Human Object Interactions

CVPR 2022poster

Modelling interactions between humans and objects in natural environments is central to many applications including gaming, virtual and mixed reality, as well as human behavior analysis and human-robot collaboration. This challenging operation scenario requires generalization to vast number of objec…

Cited by 213PDFcodeScholar
2022

Differentiable Dynamics for Articulated 3D Human Motion Reconstruction

CVPR 2022poster

We introduce DiffPhy, a differentiable physics-based model for articulated 3d human motion reconstruction from video. Applications of physics-based reasoning in human motion analysis have so far been limited, both by the complexity of constructing adequate physical models of articulated human motion…

Cited by 52PDFScholar
2022

HUM3DIL: Semi-supervised Multi-modal 3D HumanPose Estimation for Autonomous Driving

CoRL 2022poster

Autonomous driving is an exciting new industry, posing important research questions. Within the perception module, 3D human pose estimation is an emerging technology, which can enable the autonomous vehicle to perceive and understand the subtle and complex behaviors of pedestrians. While hardware sy…

Cited by 32SourceScholar
2022

Learning Online Multi-sensor Depth Fusion

ECCV 2022poster

"Many hand-held or mixed reality devices are used with a single sensor for 3D reconstruction, although they often comprise multiple sensors. Multi-sensor depth fusion is able to substantially improve the robustness and accuracy of 3D reconstruction methods, but existing techniques are not robust eno…

2022

Photorealistic Monocular 3D Reconstruction of Humans Wearing Clothing

CVPR 2022poster

We present PHORHUM, a novel, end-to-end trainable, deep neural network methodology for photorealistic 3D human reconstruction given just a monocular RGB image. Our pixel-aligned method estimates detailed 3D geometry and, for the first time, the unshaded surface color together with the scene illumina…

Cited by 161PDFScholar
2022

Trajectory Optimization for Physics-Based Reconstruction of 3D Human Pose From Monocular Video

CVPR 2022poster

We focus on the task of estimating a physically plausible articulated human motion from monocular video. Existing approaches that do not consider physics often produce temporally inconsistent output with motion artifacts, while state-of-the-art physics-based approaches have either been shown to work…

Cited by 48PDFScholar
2021

AIFit: Automatic 3D Human-Interpretable Feedback Models for Fitness Training

CVPR 2021poster

I went to the gym today, but how well did I do? And where should I improve? Ah, my back hurts slightly... User engagement can be sustained and injuries avoided by being able to reconstruct 3d human pose and motion, relate it to good training practices, identify errors, and provide early, real-time f…

Cited by 92PDFScholar
2021

Calibration of Neural Networks using Splines

ICLR 2021poster

Calibrating neural networks is of utmost importance when employing them in safety-critical applications where the downstream decision making depends on the predicted probabilities. Measuring calibration error amounts to comparing two empirical distributions. In this work, we introduce a binning-free…

2021

Embodied Visual Active Learning for Semantic Segmentation

AAAI 2021technical

We study the task of embodied visual active learning, where an agent is set to explore a 3d environment with the goal to acquire visual scene understanding by actively selecting views for which to request annotation. While accurate on some benchmarks, today's deep visual recognition pipelines tend t…

Cited by 43SourcePDFScholar
2021

Generating Scenarios with Diverse Pedestrian Behaviors for Autonomous Vehicle Testing

CoRL 2021poster

There exist several datasets for developing self-driving car methodologies. Manually collected datasets impose inherent limitations on the variability of test cases and it is particularly difficult to acquire challenging scenarios, e.g. ones involving collisions with pedestrians. A way to alleviate…

Cited by 16SourceScholar
2021

H-NeRF: Neural Radiance Fields for Rendering and Temporal Reconstruction of Humans in Motion

NeurIPS 2021spotlight

We present neural radiance fields for rendering and temporal (4D) reconstruction of humans in motion (H-NeRF), as captured by a sparse set of cameras or even from a monocular video. Our approach combines ideas from neural scene representation, novel-view synthesis, and implicit statistical geometric…

Cited by 205SourcePDFScholar
2021

Learning Complex 3D Human Self-Contact

AAAI 2021technical

Monocular estimation of three dimensional human self-contact is fundamental for detailed scene analysis including body language understanding and behaviour modeling. Existing 3d reconstruction methods do not focus on body regions in self-contact and consequently recover configurations that are eith…

Cited by 33SourcePDFScholar
2021

Neural Descent for Visual 3D Human Pose and Shape

CVPR 2021poster

We present deep neural network methodology to reconstruct the 3d pose and shape of people, including hand gestures and facial expression, given an input RGB image. We rely on a recently introduced, expressive full body statistical 3d human model, GHUM, trained end-to-end, and learn to reconstruct it…

Cited by 75PDFScholar
2021

REMIPS: Physically Consistent 3D Reconstruction of Multiple Interacting People under Weak Supervision

NeurIPS 2021poster

The three-dimensional reconstruction of multiple interacting humans given a monocular image is crucial for the general task of scene understanding, as capturing the subtleties of interaction is often the very reason for taking a picture. Current 3D human reconstruction methods either treat each pers…

Cited by 30SourcePDFScholar
2021

RSN: Range Sparse Net for Efficient, Accurate LiDAR 3D Object Detection

CVPR 2021poster

The detection of 3D objects from LiDAR data is a critical component in most autonomous driving systems. Safe, high speed driving needs larger detection ranges, which are enabled by new LiDARs. These larger detection ranges require more efficient and accurate detection models. Towards this goal, we p…

Cited by 204PDFScholar
2021

Relative Flatness and Generalization

NeurIPS 2021poster

Flatness of the loss curve is conjectured to be connected to the generalization ability of machine learning models, in particular neural networks. While it has been empirically observed that flatness measures consistently correlate strongly with generalization, it is still an open theoretical proble…

2021

THUNDR: Transformer-Based 3D Human Reconstruction With Markers

ICCV 2021poster

We present THUNDR, a transformer-based deep neural network methodology to reconstruct the 3d pose and shape of people, given monocular RGB images. Key to our methodology is an intermediate 3d marker representation, where we aim to combine the predictive power of model-free-output architectures and t…

Cited by 83PDFScholar
2021

TropEx: An Algorithm for Extracting Linear Terms in Deep Neural Networks

ICLR 2021poster

Deep neural networks with rectified linear (ReLU) activations are piecewise linear functions, where hyperplanes partition the input space into an astronomically high number of linear regions. Previous work focused on counting linear regions to measure the network's expressive power and on analyzing…

Cited by 14SourcePDFScholar
2021

imGHUM: Implicit Generative Models of 3D Human Shape and Articulated Pose

ICCV 2021poster

We present imGHUM, the first holistic generative model of 3D human shape and articulated pose, represented as a signed distance function. In contrast to prior work, we model the full human body implicitly as a function zero-level-set and without the use of an explicit template mesh. We propose a nov…

Cited by 121PDFcodeScholar
2020

Combining Implicit Function Learning and Parametric Models for 3D Human Reconstruction

ECCV 2020poster

Implicit functions represented as deep learning approximations are powerful for reconstructing 3D surfaces. However, they can only produce static surfaces that are not controllable, which provides limited ability to modify the resulting model by editing its pose or shape parameters.Implicit function…

Cited by 236SourcePDFScholar
2020

GHUM & GHUML: Generative 3D Human Shape and Articulated Pose Models

CVPR 2020oral

We present a statistical, articulated 3D human shape modeling pipeline, within a fully trainable, modular, deep learning framework. Given high-resolution complete 3D body scans of humans, captured in various poses, together with additional closeups of their head and facial expressions, as well as ha…

Cited by 423PDFcodeScholar
2020

LoopReg: Self-supervised Learning of Implicit Surface Correspondences, Pose and Shape for 3D Human Mesh Registration

NeurIPS 2020oral

We address the problem of fitting 3D human models to 3D scans of dressed humans. Classical methods optimize both the data-to-model correspondences and the human model parameters (pose and shape), but are reliable only when initialised close to the solution. Some methods initialize the optimization b…

2020

Range Conditioned Dilated Convolutions for Scale Invariant 3D Object Detection

CoRL 2020

This paper presents a novel 3D object detection framework that processes LiDAR data directly on its native representation: range images. Benefiting from the compactness of range images, 2D convolutions can efficiently process dense LiDAR data of the scene. To overcome scale sensitivity in this persp

2020

Three-Dimensional Reconstruction of Human Interactions

CVPR 2020poster

Understanding 3d human interactions is fundamental for fine grained scene analysis and behavioural modeling. However, most of the existing models focus on analyzing a single person in isolation, and those who process several people focus largely on resolving multi-person data association, rather tha…

Cited by 131PDFScholar
2020

Weakly Supervised 3D Human Pose and Shape Reconstruction with Normalizing Flows

ECCV 2020poster

Monocular 3D human pose and shape estimation is challenging due to the many degrees of freedom of the human body and the difficulty to acquire training data for large-scale supervised learning in complex visual scenes where humans with diverse shape and appearance, appear against complex backgrounds…

Cited by 160SourcePDFScholar
2019

Domes to Drones: Self-Supervised Active Triangulation for 3D Human Pose Reconstruction

NeurIPS 2019poster

Existing state-of-the-art estimation systems can detect 2d poses of multiple people in images quite reliably. In contrast, 3d pose estimation from a single image is ill-posed due to occlusion and depth ambiguities. Assuming access to multiple cameras, or given an active system able to position itsel…

2019

Self-Supervised Learning With Geometric Constraints in Monocular Video: Connecting Flow, Depth, and Camera

ICCV 2019poster

We present GLNet, a self-supervised framework for learning depth, optical flow, camera pose and intrinsic parameters from monocular video -- addressing the difficulty of acquiring realistic ground-truth for such tasks. We propose three contributions: 1) we design new loss functions that capture mult…

Cited by 317PDFScholar
2018

3D Human Sensing, Action and Emotion Recognition in Robot Assisted Therapy of Children With Autism

CVPR 2018poster

We introduce new, fine-grained action and emotion recognition tasks defined on non-staged videos, recorded during robot-assisted therapy sessions of children with autism. The tasks present several challenges: a large dataset with long videos, a large number of highly variable actions, children that…

Cited by 137SourcePDFScholar
2018

Deep Network for the Integrated 3D Sensing of Multiple People in Natural Images

NeurIPS 2018spotlight

We present MubyNet -- a feed-forward, multitask, bottom up system for the integrated localization, as well as 3d pose and shape estimation, of multiple people in monocular images. The challenge is the formal modeling of the problem that intrinsically requires discrete and continuous computation, e.g…

Cited by 170SourcePDFScholar
2018

Deep Reinforcement Learning of Region Proposal Networks for Object Detection

CVPR 2018poster

We propose drl-RPN, a deep reinforcement learning-based visual recognition model consisting of a sequential region proposal network (RPN) and an object detector. In contrast to typical RPNs, where candidate object regions (RoIs) are selected greedily via class-agnostic NMS, drl-RPN optimizes an obje…

2018

Monocular 3D Pose and Shape Estimation of Multiple People in Natural Scenes - The Importance of Multiple Scene Constraints

CVPR 2018poster

Human sensing has greatly benefited from recent advances in deep learning, parametric human modeling, and large scale 2d and 3d datasets. However, existing 3d models make strong assumptions about the scene, considering either a single person per image, full views of the person, a simple background o…

Cited by 346SourcePDFScholar
2015

Second-Order Constrained Parametric Proposals and Sequential Search-Based Structured Prediction for Semantic Segmentation in RGB-D Images

CVPR 2015poster

We focus on the problem of semantic segmentation based on RGB-D data, with emphasis on analyzing cluttered indoor scenes containing many visual categories and instances. Our approach is based on a parametric figure-ground intensity and depth-constrained proposal process that generates spatial layout…

Cited by 64SourcePDFScholar