← Search

Helge Rhodin

26 accepted papers

2026

E-3DPSM: A State Machine for Event-based Egocentric 3D Human Pose Estimation

CVPR 2026

Event cameras offer multiple advantages in monocular egocentric 3D human pose estimation from head-mounted devices, such as millisecond temporal resolution, high dynamic range, and negligible motion blur. Existing methods effectively leverage these properties, but suffer from low 3D estimation accur

Cited by 0SourceScholar
2025

Locality Sensitive Avatars From Video

ICLR 2025poster

We present locality-sensitive avatar, a neural radiance field (NeRF) based network to learn human motions from monocular videos. To this end, we estimate a canonical representation between different frames of a video with a non-linear mapping from observation to canonical space, which we decompose i…

2024

Unsupervised Keypoints from Pretrained Diffusion Models

CVPR 2024highlight

Unsupervised learning of keypoints and landmarks has seen significant progress with the help of modern neural network architectures but performance is yet to match the supervised counterpart making their practicability questionable. We leverage the emergent knowledge within text-to-image diffusion m…

2024

VIRL: Self-Supervised Visual Graph Inverse Reinforcement Learning

CoRL 2024poster

Learning dense reward functions from unlabeled videos for reinforcement learning exhibits scalability due to the vast diversity and quantity of video resources. Recent works use visual features or graph abstractions in videos to measure task progress as rewards, which either deteriorate in unseen do…

Cited by 0SourceScholar
2023

Few-Shot Geometry-Aware Keypoint Localization

CVPR 2023poster

Supervised keypoint localization methods rely on large manually labeled image datasets, where objects can deform, articulate, or occlude. However, creating such large keypoint labels is time-consuming and costly, and is often error-prone due to inconsistent labeling. Thus, we desire an approach that…

2022

AdaptPose: Cross-Dataset Adaptation for 3D Human Pose Estimation by Learnable Motion Generation

CVPR 2022poster

This paper addresses the problem of cross-dataset generalization of 3D human pose estimation models. Testing a pre-trained 3D pose estimator on a new dataset results in a major performance drop. Previous methods have mainly addressed this problem by improving the diversity of the training data. We a…

Cited by 51PDFcodeScholar
2022

AutoLink: Self-supervised Learning of Human Skeletons and Object Outlines by Linking Keypoints

NeurIPS 2022accept

Structured representations such as keypoints are widely used in pose transfer, conditional image generation, animation, and 3D reconstruction. However, their supervised learning requires expensive annotation for each target domain. We propose a self-supervised method that learns to disentangle objec…

2022

DANBO: Disentangled Articulated Neural Body Representations via Graph Neural Networks

ECCV 2022poster

"Deep learning greatly improved the realism of animatable human models by learning geometry and appearance from collections of 3D scans, template meshes, and multi-view imagery. High-resolution models enable photo-realistic avatars but at the cost of requiring studio settings not available to end us…

2022

Domain Knowledge-Informed Self-Supervised Representations for Workout Form Assessment

ECCV 2022poster

"Maintaining proper form while exercising is important for preventing injuries and maximizing muscle mass gains. Detecting errors in workout form naturally requires estimating human’s body pose. However, off-the-shelf pose estimators struggle to perform well on the videos recorded in gym scenarios d…

2022

ElePose: Unsupervised 3D Human Pose Estimation by Predicting Camera Elevation and Learning Normalizing Flows on 2D Poses

CVPR 2022poster

Human pose estimation from single images is a challenging problem that is typically solved by supervised learning. Unfortunately, labeled training data does not yet exist for many human activities since 3D annotation requires dedicated motion capture systems. Therefore, we propose an unsupervised ap…

Cited by 57PDFcodeScholar
2021

A-NeRF: Articulated Neural Radiance Fields for Learning Human Shape, Appearance, and Pose

NeurIPS 2021poster

While deep learning reshaped the classical motion capture pipeline with feed-forward networks, generative models are required to recover fine alignment via iterative refinement. Unfortunately, the existing models are usually hand-crafted or learned in controlled conditions, only applicable to limite…

Cited by 348SourcePDFScholar
2021

CanonPose: Self-Supervised Monocular 3D Human Pose Estimation in the Wild

CVPR 2021poster

Human pose estimation from single images is a challenging problem in computer vision that requires large amounts of labeled training data to be solved accurately. Unfortunately, for many human activities (e.g. outdoor sports) such training data does not exist and is hard or even impossible to acquir…

Cited by 143PDFcodeScholar
2021

Human Detection and Segmentation via Multi-View Consensus

ICCV 2021poster

Self-supervised detection and segmentation of foreground objects aims for accuracy without annotated training data. However, existing approaches predominantly rely on restrictive assumptions on appearance and motion. For scenes with dynamic activities and camera motion, we propose a multi-camera fra…

Cited by 3PDFcodeScholar
2021

PCLs: Geometry-Aware Neural Reconstruction of 3D Pose With Perspective Crop Layers

CVPR 2021poster

Local processing is an essential feature of CNNs and other neural network architectures -- it is one of the reasons why they work so well on images where relevant information is, to a large extent, local. However, perspective effects stemming from the projection in a conventional camera vary for dif…

Cited by 24PDFcodeScholar
2020

ActiveMoCap: Optimized Viewpoint Selection for Active Human Motion Capture

CVPR 2020oral

The accuracy of monocular 3D human pose estimation depends on the viewpoint from which the image is captured. While freely moving cameras, such as on drones, provide control over this viewpoint, automatically positioning them at the location which will yield the highest accuracy remains an open prob…

Cited by 47PDFcodeScholar
2020

Deformation-Aware Unpaired Image Translation for Pose Estimation on Laboratory Animals

CVPR 2020poster

Our goal is to capture the pose of real animals using synthetic training examples, without using any manual supervision. Our focus is on neuroscience model organisms, to be able to study how neural circuits orchestrate behaviour. Human pose estimation attains remarkable accuracy when trained on real…

Cited by 51PDFScholar
2020

Front2Back: Single View 3D Shape Reconstruction via Front to Back Prediction

CVPR 2020poster

Reconstruction of a 3D shape from a single 2D image is a classical computer vision problem, whose difficulty stems from the inherent ambiguity of recovering occluded or only partially observed surfaces. Recent methods address this challenge through the use of largely unstructured neural networks tha…

Cited by 51PDFcodeScholar
2019

Neural Scene Decomposition for Multi-Person Motion Capture

CVPR 2019poster

Learning general image representations has proven key to the success of many computer vision tasks. For example, many approaches to image understanding problems rely on deep networks that were initially trained on ImageNet, mostly because the learned features are a valuable starting point to learn f…

Cited by 61PDFScholar
2018

Learning Monocular 3D Human Pose Estimation From Multi-View Images

CVPR 2018poster

Accurate 3D human pose estimation from single images is possible with sophisticated deep-net architectures that have been trained on very large datasets. However, this still leaves open the problem of capturing motions for which no such database exists. Manual annotation is tedious, slow, and error…

Cited by 308SourcePDFScholar
2018

Unsupervised Geometry-Aware Representation for 3D Human Pose Estimation

ECCV 2018poster

Modern 3D human pose estimation techniques rely on deep networks, which require large amounts of training data. While weakly-supervised methods require less supervision, by utilizing 2D poses or multi-view imagery without annotations, they still need a sufficiently large set of samples with 3D annot…

Cited by 313SourcePDFScholar
2015

A Versatile Scene Model With Differentiable Visibility Applied to Generative Pose Estimation

ICCV 2015poster

Generative reconstruction methods compute the 3D configuration (such as pose and/or geometry) of a shape by optimizing the overlap of the projected 3D shape model with images. Proper handling of occlusions is a big challenge, since the visibility function that indicates if a surface point is seen fr…

Cited by 113PDFScholar