← Search

Christian Theobalt

142 accepted papers

2026

$\boldsymbol{\partial^\infty}$-Grid: A Neural Differential Equation Solver with Differentiable Feature Grids

ICLR 2026poster

We present a novel differentiable grid-based representation for efficiently solving differential equations (DEs). Widely used architectures for neural solvers, such as sinusoidal neural networks, are coordinate-based MLPs that are, both, computationally intensive and slow to train. Although grid-bas…

Cited by 0SourcecodeScholar
2026

E-3DPSM: A State Machine for Event-based Egocentric 3D Human Pose Estimation

CVPR 2026

Event cameras offer multiple advantages in monocular egocentric 3D human pose estimation from head-mounted devices, such as millisecond temporal resolution, high dynamic range, and negligible motion blur. Existing methods effectively leverage these properties, but suffer from low 3D estimation accur

Cited by 0SourceScholar
2026

EmbodMocap: In-the-Wild 4D Human-Scene Reconstruction for Embodied Agents

CVPR 2026

Human behaviors in the real world naturally encode rich, long-term contextual information that can be leveraged to train embodied agents for perception, understanding, and acting.However, existing capture systems typically rely on costly studio setups and wearable devices, limiting the large-scale c

Cited by 0SourcecodeScholar
2026

FUN REC * Reconstructing Functional 3D Scenes from Egocentric Interaction Videos

CVPR 2026

We present FunREC, a method for reconstructing functional 3D digital twins of indoor scenes directly from egocentric RGB-D interaction videos. Unlike existing methods on articulated reconstruction, which rely on controlled setups, multi-state captures, or CAD priors, FunREC operates directly on in-t

Cited by 0SourcecodeScholar
2026

MIBURI: Towards Expressive Interactive Gesture Synthesis

CVPR 2026

Embodied Conversational Agents (ECAs) aim to emulate human face-to-face interaction through speech, gestures, and facial expressions. Current large language model (LLM)-based conversational agents lack embodiment and the expressive gestures essential for natural interaction. Existing solutions for E

Cited by 0SourcecodeScholar
2026

OLATverse: A Large-scale Real-world Object Dataset with Precise Lighting Control

CVPR 2026

We introduce OLATverse, a large-scale dataset comprising around 9M images of 765 real-world objects, captured from multiple viewpoints under a diverse set of precisely controlled lighting conditions. While recent advances in object-centric inverse rendering, novel view synthesis and relighting have

Cited by 0SourcecodeScholar
2026

Physical Simulator In-the-Loop Video Generation

CVPR 2026

Recent advances in diffusion-based video generation have achieved remarkable visual realism but still struggle to obey basic physical laws such as gravity, inertia, and collision. Generated objects often move inconsistently across frames, exhibit implausible dynamics, or violate physical constraints

Cited by 0SourcecodeScholar
2026

RelightAnyone: A Generalized Relightable 3D Gaussian Head Model

CVPR 2026

3D Gaussian Splatting (3DGS) has become a standard approach to reconstruct and render photorealistic 3D head avatars. A major challenge is to relight the avatars to match any scene illumination. For high quality relighting, existing methods require subjects to be captured under complex time-multiple

Cited by 0SourceScholar
2026

Relightable Holoported Characters: Capturing and Relighting Dynamic Human Performance from Sparse Views

CVPR 2026

We present _Relightable Holoported Characters_ (RHC), a novel person-specific method for free-view rendering and relighting of full-body and highly dynamic humans solely observed from sparse-view RGB videos at inference. In contrast to classical one-light-at-a-time (OLAT)-based human relighting, our

Cited by 0SourceScholar
2026

SceMoS: Scene-Aware 3D Human Motion Synthesis by Planning with Geometry-Grounded Tokens

CVPR 2026

Synthesizing text-driven 3D human motion within realistic scenes requires learning both semantic intent ("walk to the couch") and physical feasibility (e.g., avoiding collisions). Current methods use generative frameworks that simultaneously learn high-level planning and low-level contact reasoning,

Cited by 0SourcecodeScholar
2026

Splat the Net: Radiance Fields with Splattable Neural Primitives

ICLR 2026poster

Radiance fields have emerged as a predominant representation for modeling 3D scene appearance. Neural formulations such as Neural Radiance Fields provide high expressivity but require costly ray marching for rendering, whereas primitive-based methods such as 3D Gaussian Splatting offer real-time eff…

Cited by 0SourcecodeScholar
2025

Attention (as Discrete-Time Markov) Chains

NeurIPS 2025poster

We introduce a new interpretation of the attention matrix as a discrete-time Markov chain. Our interpretation sheds light on common operations involving attention scores such as selection, summation, and averaging in a unified framework. It further extends them by considering indirect attention, pro…

Cited by 0SourceScholar
2025

BimArt: A Unified Approach for the Synthesis of 3D Bimanual Interaction with Articulated Objects

CVPR 2025poster

We present BimArt, a novel generative approach for synthesizing 3D bimanual hand interactions with articulated objects. Unlike prior works, we do not rely on a reference grasp, a coarse hand trajectory, or separate modes for grasping and articulating. To achieve this, we first generate distance-base…

Cited by 2SourcePDFScholar
2025

Bring Your Rear Cameras for Egocentric 3D Human Pose Estimation

ICCV 2025poster

Egocentric 3D human pose estimation has been actively studied using cameras installed in front of a head-mounted device (HMD). While frontal placement is the optimal and the only option for some tasks, such as hand tracking, it remains unclear if the same holds for full-body tracking due to self-occ…

2025

CNS-Bench: Benchmarking Image Classifier Robustness Under Continuous Nuisance Shifts

ICCV 2025poster

An important challenge when using computer vision models in the real world is to evaluate their performance in potential out-of-distribution (OOD) scenarios. While simple synthetic corruptions are commonly applied to test OOD robustness, they often fail to capture nuisance shifts that occur in the r…

Cited by 0SourcePDFScholar
2025

Common3D: Self-Supervised Learning of 3D Morphable Models for Common Objects in Neural Feature Space

CVPR 2025poster

3D morphable models (3DMMs) are a powerful tool to represent the possible shapes and appearances of an object category. Given a single test image, 3DMMs can be used to solve various tasks, such as predicting the 3D shape, pose, semantic correspondence, and instance segmentation of an object. Unfortu…

2025

DICE: End-to-end Deformation Capture of Hand-Face Interactions from a Single Image

ICLR 2025poster

Reconstructing 3D hand-face interactions with deformations from a single image is a challenging yet crucial task with broad applications in AR, VR, and gaming. The challenges stem from self-occlusions during single-view hand-face interactions, diverse spatial relationships between hands and face, co…

2025

Do It Yourself: Learning Semantic Correspondence from Pseudo-Labels

ICCV 2025poster

Finding correspondences between semantically similar points across images and object instances is one of the everlasting challenges in computer vision. While large pre-trained vision models have recently been demonstrated as effective priors for semantic matching, they still suffer from ambiguities…

Cited by 0SourcePDFScholar
2025

EasyHOI: Unleashing the Power of Large Models for Reconstructing Hand-Object Interactions in the Wild

CVPR 2025poster

Our work aims to reconstruct hand-object interactions from a single-view image, which is a fundamental but ill-posed task.Unlike methods that reconstruct from videos, multi-view images, or predefined 3D templates, single-view reconstruction faces significant challenges due to inherent ambiguities an…

2025

Ego4o: Egocentric Human Motion Capture and Understanding from Multi-Modal Input

CVPR 2025poster

This work focuses on tracking and understanding human motion using consumer wearable devices, such as VR/AR headsets, smart glasses, cellphones, and smartwatches. These devices provide diverse, multi-modal sensor inputs, including egocentric images, and 1-3 sparse IMU sensors in varied combinations.…

Cited by 0SourcePDFScholar
2025

Enhancing Spoken Discourse Modeling in Language Models Using Gestural Cues

ACL 2025long

Research in linguistics shows that non-verbal cues, such as gestures, play a crucial role in spoken discourse. For example, speakers perform hand gestures to indicate topic shifts, helping listeners identify transitions in discourse. In this work, we investigate whether the joint modeling of gesture…

Cited by 0SourcePDFScholar
2025

FRAME: Floor-aligned Representation for Avatar Motion from Egocentric Video

CVPR 2025highlight

Egocentric motion capture with a head-mounted body-facing stereo camera is crucial for VR and AR applications but presents significant challenges such as heavy occlusions and limited annotated real-world data. Existing methods rely on synthetic pretraining and struggle to generate smooth and accurat…

Cited by 0SourcePDFScholar
2025

HumanOLAT: A Large-Scale Dataset for Full-Body Human Relighting and Novel-View Synthesis

ICCV 2025poster

Simultaneous relighting and novel-view rendering of digital human representations is an important yet challenging task with numerous applications. However, progress in this area has been significantly limited due to the lack of publicly available, high-quality datasets, especially for full-body huma…

Cited by 0SourcePDFScholar
2025

OnlineSplatter: Pose-Free Online 3D Reconstruction for Free-Moving Objects

NeurIPS 2025spotlight

Free-moving object reconstruction from monocular video remains challenging, particularly without reliable pose or depth cues and under arbitrary object motion. We introduce OnlineSplatter, a novel online feed-forward framework generating high-quality, object-centric 3D Gaussians directly from RGB fr…

Cited by 0SourceScholar
2025

Real-time Free-view Human Rendering from Sparse-view RGB Videos using Double Unprojected Textures

CVPR 2025highlight

Real-time free-view human rendering from sparse-view RGB inputs is a challenging task due to the sensor scarcity and the tight time budget. To ensure efficiency, recent methods leverage 2D CNNs operating in texture space to learn rendering primitives. However, they either jointly learn geometry and…

Cited by 2SourcePDFScholar
2025

Retrieving Semantics from the Deep: an RAG Solution for Gesture Synthesis

CVPR 2025poster

Non-verbal communication often comprises of semantically rich gestures that help convey the meaning of an utterance. Producing such semantic co-speech gestures has been a major challenge for the existing neural systems that can generate rhythmic beat gestures, but struggle to produce semantically me…

Cited by 1SourcePDFScholar
2025

Thin-Shell-SfT: Fine-Grained Monocular Non-rigid 3D Surface Tracking with Neural Deformation Fields

CVPR 2025poster

3D reconstruction of highly deformable surfaces (e.g. cloths) from monocular RGB videos is a challenging problem, and no solution provides a consistent and accurate recovery of fine-grained surface details. To account for the ill-posed nature of the setting, existing methods use deformation models w…

Cited by 0SourcePDFScholar
2025

VidSeg: Training-free Video Semantic Segmentation based on Diffusion Models

CVPR 2025poster

We introduce the first training-free approach for Video Semantic Segmentation (VSS) based on pre-trained diffusion models. A growing research direction attempts to employ diffusion models to perform downstream vision tasks by exploiting their deep understanding of image semantics. Yet, the majority…

Cited by 0SourcePDFScholar
2024

3D Human Pose Perception from Egocentric Stereo Videos

CVPR 2024highlight

While head-mounted devices are becoming more compact they provide egocentric views with significant self-occlusions of the device user. Hence existing methods often fail to accurately estimate complex 3D poses from egocentric views. In this work we propose a new transformer-based framework to improv…

Cited by 19SourcePDFScholar
2024

ASH: Animatable Gaussian Splats for Efficient and Photoreal Human Rendering

CVPR 2024poster

Real-time rendering of photorealistic and controllable human avatars stands as a cornerstone in Computer Vision and Graphics. While recent advances in neural implicit rendering have unlocked unprecedented photorealism for digital avatars real-time performance has mostly been demonstrated for static…

2024

ConvoFusion: Multi-Modal Conversational Diffusion for Co-Speech Gesture Synthesis

CVPR 2024poster

Gestures play a key role in human communication. Recent methods for co-speech gesture generation while managing to generate beat-aligned motions struggle generating gestures that are semantically aligned with the utterance. Compared to beat gestures that align naturally to the audio signal semantica…

Cited by 12SourcePDFScholar
2024

DatasetNeRF: Efficient 3D-aware Data Factory with Generative Radiance Fields

ECCV 2024poster

"Progress in 3D computer vision tasks demands a huge amount of data, yet annotating multi-view images with 3D-consistent annotations, or point clouds with part segmentation is both time-consuming and challenging. This paper introduces DatasetNeRF, a novel approach capable of generating infinite, hig…

2024

Egocentric Whole-Body Motion Capture with FisheyeViT and Diffusion-Based Motion Refinement

CVPR 2024poster

In this work we explore egocentric whole-body motion capture using a single fisheye camera which simultaneously estimates human body and hand motion. This task presents significant challenges due to three factors: the lack of high-quality datasets fisheye camera distortion and human body self-occlus…

Cited by 21SourcePDFScholar
2024

EventEgo3D: 3D Human Motion Capture from Egocentric Event Streams

CVPR 2024poster

Monocular egocentric 3D human motion capture is a challenging and actively researched problem. Existing methods use synchronously operating visual sensors (e.g. RGB cameras) and often fail under low lighting and fast motions which can be restricting in many applications involving head-mounted device…

2024

Holoported Characters: Real-time Free-viewpoint Rendering of Humans from Sparse RGB Cameras

CVPR 2024poster

We present the first approach to render highly realistic free-viewpoint videos of a human actor in general apparel from sparse multi-view recording to display in real-time at an unprecedented 4K resolution. At inference our method only requires four camera views of the moving actor and the respectiv…

Cited by 9SourcePDFScholar
2024

MetaCap: Meta-learning Priors from Multi-View Imagery for Sparse-view Human Performance Capture and Rendering

ECCV 2024poster

"Faithful human performance capture and free-view rendering from sparse RGB observations is a long-standing problem in Vision and Graphics. The main challenges are the lack of observations and the inherent ambiguities of the setting, e.g. occlusions and depth ambiguity. As a result, radiance fields,…

Cited by 9SourcePDFScholar
2024

NeuralClothSim: Neural Deformation Fields Meet the Thin Shell Theory

NeurIPS 2024poster

Despite existing 3D cloth simulators producing realistic results, they predominantly operate on discrete surface representations (e.g. points and meshes) with a fixed spatial resolution, which often leads to large memory consumption and resolution-dependent simulations. Moreover, back-propagating gr…

2024

ReMoS: 3D Motion-Conditioned Reaction Synthesis for Two-Person Interactions

ECCV 2024poster

"Current approaches for 3D human motion synthesis generate high-quality animations of digital humans performing a wide variety of actions and gestures. However, a notable technological gap exists in addressing the complex dynamics of multi-human interactions within this paradigm. In this work, we pr…

2024

Relightable Neural Actor with Intrinsic Decomposition and Pose Control

ECCV 2024poster

"Creating a controllable and relightable digital avatar from multi-view video with fixed illumination is a very challenging problem since humans are highly articulated, creating pose-dependent appearance effects, and skin as well as clothing require space-varying BRDF modeling. Existing works on cre…

Cited by 4SourcePDFScholar
2024

Surf-D: Generating High-Quality Surfaces of Arbitrary Topologies Using Diffusion Models

ECCV 2024poster

"We present Surf-D, a novel method for generating high-quality 3D shapes as Surfaces with arbitrary topologies using Diffusion models. Previous methods explored shape generation with different representations and they suffer from limited topologies and poor geometry details. To generate high-quality…

Cited by 1SourcePDFScholar
2024

VINECS: Video-based Neural Character Skinning

CVPR 2024poster

Rigging and skinning clothed human avatars is a challenging task and traditionally requires a lot of manual work and expertise. Recent methods addressing it either generalize across different characters or focus on capturing the dynamics of a single character observed under different pose configurat…

Cited by 3SourcePDFScholar
2024

Wonder3D: Single Image to 3D using Cross-Domain Diffusion

CVPR 2024highlight

In this work we introduce Wonder3D a novel method for generating high-fidelity textured meshes from single-view images with remarkable efficiency. Recent methods based on the Score Distillation Sampling (SDS) loss methods have shown the potential to recover 3D geometry from 2D diffusion priors but t…

Cited by 414SourcePDFScholar
2023

Batch-based Model Registration for Fast 3D Sherd Reconstruction

ICCV 2023poster

3D reconstruction techniques have widely been used for digital documentation of archaeological fragments. However, efficient digital capture of fragments remains as a challenge. In this work, we aim to develop a portable, high-throughput, and accurate reconstruction system for efficient digitization…

Cited by 2PDFScholar
2023

CCuantuMM: Cycle-Consistent Quantum-Hybrid Matching of Multiple Shapes

CVPR 2023poster

Jointly matching multiple, non-rigidly deformed 3D shapes is a challenging, NP-hard problem. A perfect matching is necessarily cycle-consistent: Following the pairwise point correspondences along several shapes must end up at the starting vertex of the original shape. Unfortunately, existing quantum…

Cited by 15SourcePDFScholar
2023

DELIFFAS: Deformable Light Fields for Fast Avatar Synthesis

NeurIPS 2023poster

Generating controllable and photorealistic digital human avatars is a long-standing and important problem in Vision and Graphics. Recent methods have shown great progress in terms of either photorealism or inference speed while the combination of the two desired properties still remains unsolved. To…

Cited by 33SourcePDFScholar
2023

EventNeRF: Neural Radiance Fields From a Single Colour Event Camera

CVPR 2023poster

Asynchronously operating event cameras find many applications due to their high dynamic range, vanishingly low motion blur, low latency and low data bandwidth. The field saw remarkable progress during the last few years, and existing event-based 3D reconstruction approaches recover sparse point clou…

2023

F2-NeRF: Fast Neural Radiance Field Training With Free Camera Trajectories

CVPR 2023highlight

This paper presents a novel grid-based NeRF called F^2-NeRF (Fast-Free-NeRF) for novel view synthesis, which enables arbitrary input camera trajectories and only costs a few minutes for training. Existing fast grid-based NeRF training frameworks, like Instant-NGP, Plenoxels, DVGO, or TensoRF, are ma…

2023

GlowGAN: Unsupervised Learning of HDR Images from LDR Images in the Wild

ICCV 2023poster

Most in-the-wild images are stored in Low Dynamic Range (LDR) form, serving as a partial observation of the High Dynamic Range (HDR) visual world. Despite limited dynamic range, these LDR images are often captured with different exposures, implicitly containing information about the underlying HDR i…

Cited by 13PDFScholar
2023

Grid-Guided Neural Radiance Fields for Large Urban Scenes

CVPR 2023poster

Purely MLP-based neural radiance fields (NeRF-based methods) often suffer from underfitting with blurred renderings on large-scale scenes due to limited model capacity. Recent approaches propose to geographically divide the scene and adopt multiple sub-NeRFs to model each region individually, leadin…

Cited by 94SourcePDFScholar
2023

Imitator: Personalized Speech-driven 3D Facial Animation

ICCV 2023poster

Speech-driven 3D facial animation has been widely explored, with applications in gaming, character animation, virtual reality, and telepresence systems. State-of-the-art methods deform the face topology of the target actor to sync the input audio without considering the identity-specific speaking st…

Cited by 60PDFcodeScholar
2023

LiveHand: Real-time and Photorealistic Neural Hand Rendering

ICCV 2023poster

The human hand is the main medium through which we interact with our surroundings, making its digitization an important problem. While there are several works modeling the geometry of hands, little attention has been paid to capturing photo-realistic appearance. Moreover, for applications in extende…

Cited by 19PDFcodeScholar
2023

Mofusion: A Framework for Denoising-Diffusion-Based Motion Synthesis

CVPR 2023highlight

Conventional methods for human motion synthesis have either been deterministic or have had to struggle with the trade-off between motion diversity vs motion quality. In response to these limitations, we introduce MoFusion, i.e., a new denoising-diffusion-based framework for high-quality conditional…

Cited by 187SourcePDFScholar
2023

NerfDiff: Single-image View Synthesis with NeRF-guided Distillation from 3D-aware Diffusion

ICML 2023poster

Novel view synthesis from a single image requires inferring occluded regions of objects and scenes whilst simultaneously maintaining semantic and physical consistency with the input. Existing approaches condition neural radiance fields (NeRF) on local image features, projecting points to the input i…

Cited by 182SourcePDFScholar
2023

NeuS2: Fast Learning of Neural Implicit Surfaces for Multi-view Reconstruction

ICCV 2023poster

Recent methods for neural surface representation and rendering, for example NeuS, have demonstrated the remarkably high-quality reconstruction of static scenes. However, the training of NeuS takes an extremely long time (8 hours), which makes it almost impossible to apply them to dynamic scenes with…

Cited by 276PDFcodeScholar
2023

NeuralUDF: Learning Unsigned Distance Fields for Multi-View Reconstruction of Surfaces With Arbitrary Topologies

CVPR 2023poster

We present a novel method, called NeuralUDF, for reconstructing surfaces with arbitrary topologies from 2D images via volume rendering. Recent advances in neural rendering based reconstruction have achieved compelling results. However, these methods are limited to objects with closed surfaces since…

Cited by 67SourcePDFScholar
2023

Regularized Vector Quantization for Tokenized Image Synthesis

CVPR 2023poster

Quantizing images into discrete representations has been a fundamental problem in unified generative modeling. Predominant approaches learn the discrete representation either in a deterministic manner by selecting the best-matching token or in a stochastic manner by sampling from a predicted distrib…

2023

Scene-Aware Egocentric 3D Human Pose Estimation

CVPR 2023poster

Egocentric 3D human pose estimation with a single head-mounted fisheye camera has recently attracted attention due to its numerous applications in virtual and augmented reality. Existing methods still struggle in challenging poses where the human body is highly occluded or is closely interacting wit…

2023

Voxurf: Voxel-based Efficient and Accurate Neural Surface Reconstruction

ICLR 2023top-25%

Neural surface reconstruction aims to reconstruct accurate 3D surfaces based on multi-view images. Previous methods based on neural volume rendering mostly train a fully implicit model with MLPs, which typically require hours of training for a single scene. Recent efforts explore the explicit volume…

2023

WaveNeRF: Wavelet-based Generalizable Neural Radiance Fields

ICCV 2023poster

Neural Radiance Field (NeRF) has shown impressive performance in novel view synthesis via implicit scene representation. However, it usually suffers from poor scalability as requiring densely sampled images for each new scene. Several studies have attempted to mitigate this problem by integrating Mu…

Cited by 16PDFScholar
2023

Weakly Supervised 3D Open-vocabulary Segmentation

NeurIPS 2023poster

Open-vocabulary segmentation of 3D scenes is a fundamental function of human perception and thus a crucial objective in computer vision research. However, this task is heavily impeded by the lack of large-scale and diverse 3D open-vocabulary segmentation datasets for training robust and generalizabl…

2022

BEHAVE: Dataset and Method for Tracking Human Object Interactions

CVPR 2022poster

Modelling interactions between humans and objects in natural environments is central to many applications including gaming, virtual and mixed reality, as well as human behavior analysis and human-robot collaboration. This challenging operation scenario requires generalization to vast number of objec…

Cited by 213PDFcodeScholar
2022

BungeeNeRF: Progressive Neural Radiance Field for Extreme Multi-Scale Scene Rendering

ECCV 2022poster

"Neural Radiance Field (NeRF) has achieved outstanding performance in modeling 3D objects and controlled scenes, usually under a single scale. In this work, we focus on multi-scale cases where large changes in imagery are observed at drastically different scales. This scenario vastly exists in the r…

Cited by 267SourcePDFScholar
2022

Disentangled3D: Learning a 3D Generative Model With Disentangled Geometry and Appearance From Monocular Images

CVPR 2022poster

Learning 3D generative models from a dataset of monocular images enables self-supervised 3D reasoning and controllable synthesis. State-of-the-art 3D generative models are GANs which use neural 3D volumetric representations for synthesis. Images are synthesized by rendering the volumes from a given…

Cited by 51PDFScholar
2022

Estimating Egocentric 3D Human Pose in the Wild With External Weak Supervision

CVPR 2022poster

Egocentric 3D human pose estimation with a single fisheye camera has drawn a significant amount of attention recently. However, existing methods struggle with pose estimation from in-the-wild images, because they can only be trained on synthetic data due to the unavailability of large-scale in-the-w…

Cited by 37PDFScholar
2022

HULC: 3D HUman Motion Capture with Pose Manifold SampLing and Dense Contact Guidance

ECCV 2022poster

"Marker-less monocular 3D human motion capture (MoCap) with scene interactions is a challenging research topic relevant for extended reality, robotics and virtual avatar generation. Due to the inherent depth ambiguity of monocular settings, 3D motions captured with existing methods often contain sev…

Cited by 31SourcePDFScholar
2022

Learn to Predict How Humans Manipulate Large-Sized Objects From Interactive Motions

RA-L 2022

Understanding human intentions during interactions has been a long-lasting theme, that has applications in human-robot interaction, virtual reality and surveillance. In this study, we focus on full-body human interactions with large-sized daily objects and aim to predict the future states of objects

Cited by 35SourceScholar
2022

NeRF for Outdoor Scene Relighting

ECCV 2022poster

"Photorealistic editing of outdoor scenes from photographs requires a profound understanding of the image formation process and an accurate estimation of the scene geometry, reflectance and illumination. A delicate manipulation of the lighting can then be performed while keeping the scene albedo and…

Cited by 150SourcePDFScholar
2022

NeuRIS: Neural Reconstruction of Indoor Scenes Using Normal Priors

ECCV 2022poster

"Reconstructing 3D indoor scenes from 2D images is an important task in many computer vision and graphics applications. A main challenge in this task is that large texture-less areas in typical indoor scenes make existing methods struggle to produce satisfactory reconstruction results. We propose a…

Cited by 113SourcePDFScholar
2022

Neural Radiance Transfer Fields for Relightable Novel-View Synthesis with Global Illumination

ECCV 2022poster

"Given a set of images of a scene, the re-rendering of this scene from novel views and lighting conditions is an important and challenging problem in Computer Vision and Graphics. On the one hand, most existing works in Computer Vision usually impose many assumptions regarding the image formation pr…

Cited by 50SourcePDFScholar
2022

Neural Rays for Occlusion-Aware Image-Based Rendering

CVPR 2022poster

We present a new neural representation, called Neural Ray (NeuRay), for the novel view synthesis task. Recent works construct radiance fields from image features of input views to render novel view images, which enables the generalization to new scenes. However, due to occlusions, a 3D point may be…

Cited by 234PDFcodeScholar
2022

Physical Inertial Poser (PIP): Physics-Aware Real-Time Human Motion Tracking From Sparse Inertial Sensors

CVPR 2022poster

Motion capture from sparse inertial sensors has shown great potential compared to image-based approaches since occlusions do not lead to a reduced tracking quality and the recording space is not restricted to be within the viewing frustum of the camera. However, capturing the motion and global posit…

Cited by 200PDFScholar
2022

Playable Environments: Video Manipulation in Space and Time

CVPR 2022poster

We present Playable Environments - a new representation for interactive video generation and manipulation in space and time. With a single image at inference time, our novel framework allows the user to move objects in 3D while generating a video by providing a sequence of desired actions. The actio…

Cited by 20PDFcodeScholar
2022

StyleNeRF: A Style-based 3D Aware Generator for High-resolution Image Synthesis

ICLR 2022poster

We propose StyleNeRF, a 3D-aware generative model for photo-realistic high-resolution image synthesis with high multi-view consistency, which can be trained on unstructured 2D images. Existing approaches either cannot synthesize high-resolution images with fine details or yield clearly noticeable 3…

2022

UnrealEgo: A New Dataset for Robust Egocentric 3D Human Motion Capture

ECCV 2022poster

"We present UnrealEgo, a new large-scale naturalistic dataset for egocentric 3D human pose estimation. UnrealEgo is based on an advanced concept of eyeglasses equipped with two fisheye cameras that can be used in unconstrained environments. We design their virtual prototype and attach them to 3D hum…

Cited by 54SourcePDFScholar
2022

f-SfT: Shape-From-Template With a Physics-Based Deformation Model

CVPR 2022poster

Shape-from-Template (SfT) methods estimate 3D surface deformations from a single monocular RGB camera while assuming a 3D state known in advance (a template). This is an important yet challenging problem due to the under-constrained nature of the monocular setting. Existing SfT techniques predominan…

Cited by 24PDFcodeScholar
2021

A Shading-Guided Generative Implicit Model for Shape-Accurate 3D-Aware Image Synthesis

NeurIPS 2021poster

The advancement of generative radiance fields has pushed the boundary of 3D-aware image synthesis. Motivated by the observation that a 3D object should look realistic from multiple viewpoints, these methods introduce a multi-view constraint as regularization to learn valid 3D radiance fields from 2D…

2021

Adaptive Surface Normal Constraint for Depth Estimation

ICCV 2021poster

We present a novel method for single image depth estimation using surface normal constraints. Existing depth estimation methods either suffer from the lack of geometric constraints, or are limited to the difficulty of reliably capturing geometric context, which leads to a bottleneck of depth estimat…

Cited by 72PDFcodeScholar
2021

Efficient and Differentiable Shadow Computation for Inverse Problems

ICCV 2021poster

Differentiable rendering has received increasing interest in the solution of image-based inverse problems. It can benefit traditional optimization-based solutions to inverse problems, but also allows for self-supervision of learning-based approaches for which training data with ground truth annotati…

Cited by 14PDFScholar
2021

EgoRenderer: Rendering Human Avatars From Egocentric Camera Images

ICCV 2021poster

We present EgoRenderer, a system for rendering full-body neural avatars of a person captured by a wearable, egocentric fisheye camera that is mounted on a cap or a VR headset. Our system renders photorealistic novel views of the actor and her motion from arbitrary virtual camera locations. Rendering…

Cited by 16PDFScholar
2021

Estimating Egocentric 3D Human Pose in Global Space

ICCV 2021poster

Egocentric 3D human pose estimation using a single fisheye camera has become popular recently as it allows capturing a wide range of daily activities in unconstrained environments, which is difficult for traditional outside-in motion capture with external cameras. However, existing methods have seve…

Cited by 83PDFcodeScholar
2021

EventHands: Real-Time Neural 3D Hand Pose Estimation From an Event Stream

ICCV 2021poster

3D hand pose estimation from monocular videos is a long-standing and challenging problem, which is now seeing a strong upturn. In this work, we address it for the first time using a single event camera, i.e., an asynchronous vision sensor reacting on brightness changes. Our EventHands approach has c…

Cited by 62PDFcodeScholar
2021

Gravity-Aware Monocular 3D Human-Object Reconstruction

ICCV 2021poster

This paper proposes GraviCap, i.e., a new approach for joint markerless 3D human motion capture and object trajectory estimation from monocular RGB videos. We focus on scenes with objects partially observed during a free flight. In contrast to existing monocular methods, we can recover scale, object…

Cited by 31PDFScholar
2021

High-Fidelity Neural Human Motion Transfer From Monocular Video

CVPR 2021poster

Video-based human motion transfer creates video animations of humans following a source motion. Current methods show remarkable results for tightly-clad subjects. However, the lack of temporally consistent handling of plausible clothing dynamics, including fine and high-frequency details, significan…

Cited by 41PDFScholar
2021

Learning Complete 3D Morphable Face Models From Images and Videos

CVPR 2021poster

Most 3D face reconstruction methods rely on 3D morphable models, which disentangle the space of facial deformations into identity and expression geometry, and skin reflectance. These models are typically learned from a limited number of 3D scans and thus do not generalize well across different ident…

Cited by 56PDFScholar
2021

Monocular Real-Time Full Body Capture With Inter-Part Correlations

CVPR 2021poster

We present the first method for real-time full body capture that estimates shape and motion of body and hands together with a dynamic 3D face model from a single color image. Our approach uses a new neural network architecture that exploits correlations between body and hands at high computational e…

Cited by 72PDFScholar
2021

Monocular Reconstruction of Neural Face Reflectance Fields

CVPR 2021poster

The reflectance field of a face describes the reflectance properties responsible for complex lighting effects including diffuse, specular, inter-reflection and self shadowing. Most existing methods for estimating the face reflectance from a monocular image assume faces to be diffuse with very few ap…

Cited by 36PDFScholar
2021

Multi-view Depth Estimation using Epipolar Spatio-Temporal Networks

CVPR 2021poster

We present a novel method for multi-view depth estimation from a single video, which is a critical task in various applications, such as perception, reconstruction and robot navigation. Although previous learning-based methods have demonstrated compelling results, most works estimate depth maps of i…

Cited by 88PDFcodeScholar
2021

NeuS: Learning Neural Implicit Surfaces by Volume Rendering for Multi-view Reconstruction

NeurIPS 2021spotlight

We present a novel neural surface reconstruction method, called NeuS, for reconstructing objects and scenes with high fidelity from 2D image inputs. Existing neural surface reconstruction approaches, such as DVR [Niemeyer et al., 2020] and IDR [Yariv et al., 2020], require foreground mask as supervi…

2021

Non-Rigid Neural Radiance Fields: Reconstruction and Novel View Synthesis of a Dynamic Scene From Monocular Video

ICCV 2021poster

We present Non-Rigid Neural Radiance Fields (NR-NeRF), a reconstruction and novel view synthesis approach for general non-rigid dynamic scenes. Our approach takes RGB images of a dynamic scene as input (e.g., from a monocular video recording), and creates a high-quality space-time geometry and appea…

Cited by 557PDFScholar
2021

Pose-Guided Human Animation From a Single Image in the Wild

CVPR 2021poster

We present a new pose transfer method for synthesizing a human animation from a single image of a person controlled by a sequence of body poses. Existing pose transfer methods exhibit significant visual artifacts when applying to a novel scene, resulting in temporal inconsistency and failures in pre…

Cited by 77PDFScholar
2021

Q-Match: Iterative Shape Matching via Quantum Annealing

ICCV 2021poster

Finding shape correspondences can be formulated as an NP-hard quadratic assignment problem (QAP) that becomes infeasible for shapes with high sampling density. A promising research direction is to tackle such quadratic optimization problems over binary variables with quantum annealing, which allows…

Cited by 37PDFScholar
2021

Synthesis of Compositional Animations From Textual Descriptions

ICCV 2021poster

How can we animate 3D-characters from a movie script or move robots by simply telling them what we would like them to do?" How unstructured and complex can we make a sentence and still generate plausible movements from it?" These are questions that need to be answered in the long-run, as the field i…

Cited by 202PDFcodeScholar
2021

Towards High Fidelity Monocular Face Reconstruction With Rich Reflectance Using Self-Supervised Learning and Ray Tracing

ICCV 2021poster

Robust face reconstruction from monocular image in general lighting conditions is challenging. Methods combining deep neural network encoders with differentiable rendering have opened up the path for very fast monocular reconstruction of geometry, lighting and reflectance. They can also be trained i…

Cited by 56PDFcodeScholar
2021

i3DMM: Deep Implicit 3D Morphable Model of Human Heads

CVPR 2021poster

We present the first deep implicit 3D morphable model (i3DMM) of full heads. Unlike earlier morphable face models it not only captures identity-specific geometry, texture, and expressions of the frontal face, but also models the entire head, including hair. We collect a new dataset consisting of 64…

Cited by 139PDFcodeScholar
2020

Combining Implicit Function Learning and Parametric Models for 3D Human Reconstruction

ECCV 2020poster

Implicit functions represented as deep learning approximations are powerful for reconstructing 3D surfaces. However, they can only produce static surfaces that are not controllable, which provides limited ability to modify the resulting model by editing its pose or shape parameters.Implicit function…

Cited by 236SourcePDFScholar
2020

DEMEA: Deep Mesh Autoencoders for Non-Rigidly Deforming Objects

ECCV 2020poster

Mesh autoencoders are commonly used for dimensionality reduction, sampling and mesh modeling. We propose a general-purpose DEep MEsh Autoencoder \hbox{(DEMEA)} which adds a novel embedded deformation layer to a graph-convolutional mesh autoencoder. The embedded deformation layer (EDL) is a different…

Cited by 51SourcePDFScholar
2020

DeepCap: Monocular Human Performance Capture Using Weak Supervision

CVPR 2020oral

Human performance capture is a highly important computer vision problem with many applications in movie production and virtual/augmented reality. Many previous performance capture approaches either required expensive multi-view setups or did not recover dense space-time coherent geometry with frame-…

Cited by 266PDFScholar
2020

DeepDeform: Learning Non-Rigid RGB-D Reconstruction With Semi-Supervised Data

CVPR 2020poster

Applying data-driven approaches to non-rigid 3D reconstruction has been difficult, which we believe can be attributed to the lack of a large-scale training corpus. Unfortunately, this method fails for important cases such as highly non-rigid deformations. We first address this problem of lack of dat…

Cited by 102PDFcodeScholar
2020

EventCap: Monocular 3D Capture of High-Speed Human Motions Using an Event Camera

CVPR 2020oral

The high frame rate is a critical requirement for capturing fast human motions. In this setting, existing markerless image-based methods are constrained by the lighting requirement, the high data bandwidth and the consequent high computation overhead. In this paper, we propose EventCap -- the first…

Cited by 124PDFScholar
2020

HTML: A Parametric Hand Texture Model for 3D Hand Reconstruction and Personalization

ECCV 2020poster

3D hand reconstruction from images is a widely-studied problem in computer vision and graphics, and has a particularly high relevance for virtual and augmented reality. Although several 3D hand reconstruction approaches leverage hand models as a strong prior to resolve ambiguities and achieve more r…

Cited by 87SourcePDFScholar
2020

HandVoxNet: Deep Voxel-Based Network for 3D Hand Shape and Pose Estimation From a Single Depth Map

CVPR 2020poster

3D hand shape and pose estimation from a single depth map is a new and challenging computer vision problem with many applications. The state-of-the-art methods directly regress 3D hand meshes from 2D depth images via 2D convolutional neural networks, which leads to artefacts in the estimations due t…

Cited by 93PDFScholar
2020

Image-guided Neural Object Rendering

ICLR 2020poster

We propose a learned image-guided rendering technique that combines the benefits of image-based rendering and GAN-based image synthesis. The goal of our method is to generate photo-realistic re-renderings of reconstructed objects for virtual and augmented reality applications (e.g., virtual showroom…

Cited by 71SourceScholar
2020

LoopReg: Self-supervised Learning of Implicit Surface Correspondences, Pose and Shape for 3D Human Mesh Registration

NeurIPS 2020oral

We address the problem of fitting 3D human models to 3D scans of dressed humans. Classical methods optimize both the data-to-model correspondences and the human model parameters (pose and shape), but are reliable only when initialised close to the solution. Some methods initialize the optimization b…

2020

MINA: Convex Mixed-Integer Programming for Non-Rigid Shape Alignment

CVPR 2020poster

We present a convex mixed-integer programming formulation for non-rigid shape matching. To this end, we propose a novel shape deformation model based on an efficient low-dimensional discrete model, so that finding a globally optimal solution is tractable in (most) practical cases. Our approach combi…

Cited by 25PDFScholar
2020

Monocular Real-Time Hand Shape and Motion Capture Using Multi-Modal Data

CVPR 2020poster

We present a novel method for monocular hand shape and pose estimation at unprecedented runtime performance of 100fps and at state-of-the-art accuracy. This is enabled by a new learning based architecture designed such that it can make use of all the sources of available hand training data: image da…

Cited by 257PDFcodeScholar
2020

Neural Dense Non-Rigid Structure from Motion with Latent Space Constraints

ECCV 2020poster

We introduce the first dense neural non-rigid structure from motion (N-NRSfM) approach, which can be trained end-to-end in an unsupervised manner from 2D point tracks. Compared to the competing methods, our combination of loss functions is fully-differentiable and can be readily integrated into deep…

Cited by 66SourcePDFScholar
2020

Neural Re-Rendering of Humans from a Single Image

ECCV 2020poster

Human re-rendering from a single image is a starkly under-constrained problem and state-of-the-art algorithms often exhibit un-desired artefacts, such as oversmoothing, unrealistic distortions of thebody parts and garments, or implausible changes of the texture. To ad-dress these challenges, we prop…

Cited by 91SourcePDFScholar
2020

Neural Sparse Voxel Fields

NeurIPS 2020spotlight

Photo-realistic free-viewpoint rendering of real-world scenes using classical computer graphics techniques is challenging, because it requires the difficult step of capturing detailed appearance and geometry models. Recent studies have demonstrated promising results by learning scene representations…

2020

Neural Voice Puppetry: Audio-driven Facial Reenactment

ECCV 2020poster

We present Neural Voice Puppetry, a novel approach for audio-driven facial video synthesis. Given an audio sequence of a source person or digital assistant, we generate a photo-realistic output video of a target person that is in sync with the audio of the source input. This audio-driven facial reen…

2020

Occlusion-Aware Depth Estimation with Adaptive Normal Constraints

ECCV 2020poster

We present a new learning-based method for multi-frame depth estimation from a color video, which is a fundamental problem in scene understanding, robot navigation or handheld 3D reconstruction. While recent learning-based methods estimate depth at high accuracy, 3D point clouds exported from their…

2020

PatchNets: Patch-Based Generalizable Deep Implicit 3D Shape Representations

ECCV 2020poster

Implicit surface representation combined with deep learning has led to impressive models which can represent detailed shapes of objects. Implicit surface representations, such as signed-distance functions, allow to represent shapes of arbitrary topologies. Since a continous function is learned, the…

Cited by 116SourcePDFScholar
2020

Self-supervised Outdoor Scene Relighting

ECCV 2020poster

Outdoor scene relighting is a challenging problem that requires good understanding of the scene geometry, illumination and albedo. Current techniques are completely supervised, requiring high quality synthetic renderings to train a solution. Such renderings are synthesized using priors learned from…

Cited by 62SourcePDFScholar
2020

StyleRig: Rigging StyleGAN for 3D Control Over Portrait Images

CVPR 2020oral

StyleGAN generates photorealistic portrait images of faces with eyes, teeth, hair and context (neck, shoulders, background), but lacks a rig-like control over semantic face parameters that are interpretable in 3D, such as face pose, expressions, and scene illumination. Three-dimensional morphable fa…

Cited by 473PDFScholar
2019

A Convex Relaxation for Multi-Graph Matching

CVPR 2019oral

We present a convex relaxation for the multi-graph matching problem. Our formulation allows for partial pairwise matchings, guarantees cycle consistency, and our objective incorporates both linear and quadratic costs. Moreover, we also present an extension to higher-order costs. In order to solve th…

Cited by 55PDFcodeScholar
2019

Accelerated Gravitational Point Set Alignment With Altered Physical Laws

ICCV 2019poster

This work describes Barnes-Hut Rigid Gravitational Approach (BH-RGA) -- a new rigid point set registration method relying on principles of particle dynamics. Interpreting the inputs as two interacting particle swarms, we directly minimise the gravitational potential energy of the system using non-li…

Cited by 16PDFcodeScholar
2019

FML: Face Model Learning From Videos

CVPR 2019oral

Monocular image-based 3D reconstruction of faces is a long-standing problem in computer vision. Since image data is a 2D projection of a 3D face, the resulting depth ambiguity makes the problem ill-posed. Most existing methods rely on data-driven priors that are built from limited 3D face scans. In…

Cited by 179PDFScholar
2019

HiPPI: Higher-Order Projected Power Iterations for Scalable Multi-Matching

ICCV 2019poster

The matching of multiple objects (e.g. shapes or images) is a fundamental problem in vision and graphics. In order to robustly handle ambiguities, noise and repetitive patterns in challenging real-world settings, it is essential to take geometric consistency between points into account. Computationa…

Cited by 40PDFScholar
2019

In the Wild Human Pose Estimation Using Explicit 2D Features and Intermediate 3D Representations

CVPR 2019oral

Convolutional Neural Network based approaches for monocular 3D human pose estimation usually require a large amount of training images with 3D pose annotations. While it is feasible to provide 2D joint annotations for large corpora of in-the-wild images with humans, providing accurate 3D annotations…

Cited by 178PDFScholar
2019

Learning to Reconstruct People in Clothing From a Single RGB Camera

CVPR 2019poster

We present Octopus, a learning-based model to infer the personalized 3D shape of people from a few frames (1-8) of a monocular video in which the person is moving with a reconstruction accuracy of 4 to 5mm, while being orders of magnitude faster than previous methods. From semantic segmentation imag…

Cited by 380PDFcodeScholar
2019

Multi-Garment Net: Learning to Dress 3D People From Images

ICCV 2019poster

We present Multi-Garment Network (MGN), a method to predict body shape and clothing, layered on top of the SMPL model from a few frames (1-8) of a video. Several experiments demonstrate that this representation allows higher level of control when compared to single mesh or voxel representations of s…

Cited by 464PDFScholar
2019

Tex2Shape: Detailed Full Human Body Geometry From a Single Image

ICCV 2019poster

We present a simple yet effective method to infer detailed full human body shape from only a single photograph. Our model can infer full-body shape including face, hair, and clothing including wrinkles at interactive frame-rates. Results feature details even on parts that are occluded in the input i…

Cited by 372PDFcodeScholar
2018

A Hybrid Model for Identity Obfuscation by Face Replacement

ECCV 2018poster

As more and more personal photos are shared and tagged in social media, avoiding privacy risks such as unintended recognition, becomes increasingly challenging. We propose a new hybrid approach to obfuscate identities in photos by head replacement. Our approach combines state of the art parametric f…

Cited by 143SourcePDFScholar
2018

DS*: Tighter Lifting-Free Convex Relaxations for Quadratic Matching Problems

CVPR 2018poster

In this work we study convex relaxations of quadratic optimisation problems over permutation matrices. While existing semidefinite programming approaches can achieve remarkably tight relaxations, they have the strong disadvantage that they lift the original n^2-dimensional variable to an n^4-dimensi…

Cited by 53SourcePDFScholar
2018

GANerated Hands for Real-Time 3D Hand Tracking From Monocular RGB

CVPR 2018poster

We address the highly challenging problem of real-time 3D hand tracking based on a monocular RGB-only sequence. Our tracking method combines a convolutional neural network with a kinematic 3D hand model, such that it generalizes well to unseen data, is robust to occlusions and varying camera viewpoi…

Cited by 669SourcePDFScholar
2018

InverseFaceNet: Deep Monocular Inverse Face Rendering

CVPR 2018poster

We introduce InverseFaceNet, a deep convolutional inverse rendering framework for faces that jointly estimates facial pose, shape, expression, reflectance and illumination from a single input image. By estimating all parameters from just a single image, advanced editing possibilities on a single fac…

Cited by 76SourcePDFScholar
2018

LIME: Live Intrinsic Material Estimation

CVPR 2018poster

We present the first end-to-end approach for real-time material estimation for general object shapes with uniform material that only requires a single color image as input. In addition to Lambertian surface properties, our approach fully automatically computes the specular albedo, material shininess…

2018

Self-Supervised Multi-Level Face Model Learning for Monocular Reconstruction at Over 250 Hz

CVPR 2018poster

The reconstruction of dense 3D models of face geometry and appearance from a single image is highly challenging and ill-posed. To constrain the problem, many approaches rely on strong priors, such as parametric face models learned from limited 3D scan data. However, prior models restrict generalizat…

Cited by 308SourcePDFScholar
2018

Video Based Reconstruction of 3D People Models

CVPR 2018poster

This paper describes how to obtain accurate 3D body models and texture of arbitrary people from a single, monocular video in which a person is moving. Based on a parametric body model, we present a robust processing pipeline achieving 3D model fits with 5mm accuracy also for clothed people. Our main…

2017

MoFA: Model-Based Deep Convolutional Face Autoencoder for Unsupervised Monocular Reconstruction

ICCV 2017oral

In this work we propose a novel model-based deep convolutional autoencoder that addresses the highly challenging problem of reconstructing a 3D human face from a single in-the-wild color image. To this end, we combine a convolutional encoder network with an expert-designed generative model that serv…

Cited by 688PDFScholar
2017

Real-Time Hand Tracking Under Occlusion From an Egocentric RGB-D Sensor

ICCV 2017poster

We present an approach for real-time, robust, and accurate hand pose estimation from moving egocentric RGB-D cameras in cluttered real environments. Existing methods typically fail for hand-object interactions in cluttered scenes imaged from egocentric viewpoints, common for virtual or augmented rea…

Cited by 409PDFScholar
2016

Face2Face: Real-Time Face Capture and Reenactment of RGB Videos

CVPR 2016oral

We present a novel approach for real-time facial reenactment of a monocular target video sequence (e.g., Youtube video). The source sequence is also a monocular video stream, captured live with a commodity webcam. Our goal is to animate the facial expressions of the target video by a source actor an…

Cited by 2654PDFScholar
2015

A Versatile Scene Model With Differentiable Visibility Applied to Generative Pose Estimation

ICCV 2015poster

Generative reconstruction methods compute the 3D configuration (such as pose and/or geometry) of a shape by optimizing the overlap of the projected 3D shape model with images. Proper handling of occlusions is a big challenge, since the visibility function that indicates if a surface point is seen fr…

Cited by 113PDFScholar
2015

Context-Guided Diffusion for Label Propagation on Graphs

ICCV 2015poster

Existing approaches for diffusion on graphs, e.g., for label propagation, are mainly focused on isotropic diffusion, which is induced by the commonly-used graph Laplacian regularizer. Inspired by the success of diffusivity tensors for anisotropic diffusion in image processing, we presents anisotropi…

Cited by 21PDFScholar
2015

Efficient ConvNet-Based Marker-Less Motion Capture in General Scenes With a Low Number of Cameras

CVPR 2015poster

We present a novel method for accurate marker-less capture of articulated skeleton motion of several subjects in general scenes, indoors and outdoors, even from input filmed with as few as two cameras. Our approach unites a discriminative image-based joint detection method with a model-based generat…

Cited by 192SourcePDFScholar
2015

Fast and Robust Hand Tracking Using Detection-Guided Optimization

CVPR 2015poster

Markerless tracking of hands and fingers is a promising enabler for human-computer interaction. However, adoption has been limited because of tracking inaccuracies, incomplete coverage of motions, low framerate, complex camera setups, and high computational requirements. In this paper, we present a…

Cited by 298SourcePDFScholar
2015

Local High-Order Regularization on Data Manifolds

CVPR 2015poster

The common graph Laplacian regularizer is well-established in semi-supervised learning and spectral dimensionality reduction. However, as a first-order regularizer, it can lead to degenerate functions in high-dimensional manifolds. The iterated graph Laplacian enables high-order regularization, but…

Cited by 9SourcePDFScholar
2015

Semi-Supervised Learning With Explicit Relationship Regularization

CVPR 2015poster

In many learning tasks, the structure of the target space of a function holds rich information about the relationships between evaluations of functions on different data points. Existing approaches attempt to exploit this relationship information implicitly by enforcing smoothness on function evalua…

Cited by 11SourcePDFScholar