← Search

Cem Keskin

21 accepted papers

2026

EgoPoseFormer v2: Accurate Egocentric Human Motion Estimation for AR/VR

CVPR 2026

Egocentric 3D human motion estimation is essential for AR/VR experiences, yet remains challenging due to limited body coverage from the egocentric viewpoint, frequent occlusions, and scarce labeled data. We present EgoPoseFormer v2, a method that addresses these challenges through two key contributi

Cited by 0SourceScholar
2026

Geometric Neural Distance Fields for Learning Human Motion Priors

CVPR 2026

We introduce Neural Riemannian Motion Fields (\name), a novel 3D generative human motion prior that enables robust, temporally consistent, and physically plausible 3D motion recovery. Unlike existing VAE or diffusion-based methods, our higher-order motion prior explicitly models the human motion in

Cited by 0SourceScholar
2025

FoundHand: Large-Scale Domain-Specific Learning for Controllable Hand Image Generation

CVPR 2025highlight

Despite remarkable progress in image generation models, generating realistic hands remains a persistent challenge due to their complex articulation, varying viewpoints, and frequent occlusions. We present FoundHand, a large-scale domain-specific diffusion model for synthesizing single and dual hand…

Cited by 0SourcePDFScholar
2024

EgoPoseFormer: A Simple Baseline for Stereo Egocentric 3D Human Pose Estimation

ECCV 2024poster

"We present , a simple yet effective transformer-based model for stereo egocentric human pose estimation. The main challenge in egocentric pose estimation is overcoming joint invisibility, which is caused by self-occlusion or a limited field of view (FOV) of head-mounted cameras. Our approach overco…

2024

FoundPose: Unseen Object Pose Estimation with Foundation Features

ECCV 2024poster

"We propose FoundPose, a model-based method for 6D pose estimation of unseen objects from a single RGB image. The method can quickly onboard new objects using their 3D models without requiring any object- or task-specific training. In contrast, existing methods typically pre-train on large-scale, ta…

2023

AssemblyHands: Towards Egocentric Activity Understanding via 3D Hand Pose Estimation

CVPR 2023poster

We present AssemblyHands, a large-scale benchmark dataset with accurate 3D hand pose annotations, to facilitate the study of egocentric activities with challenging hand-object interactions. The dataset includes synchronized egocentric and exocentric images sampled from the recent Assembly101 dataset…

Cited by 67SourcePDFScholar
2023

In-Hand 3D Object Scanning From an RGB Sequence

CVPR 2023poster

We propose a method for in-hand 3D scanning of an unknown object with a monocular camera. Our method relies on a neural implicit surface representation that captures both the geometry and the appearance of the object, however, by contrast with most NeRF-based methods, we do not assume that the camer…

Cited by 24SourcePDFScholar
2023

MotionDeltaCNN: Sparse CNN Inference of Frame Differences in Moving Camera Videos with Spherical Buffers and Padded Convolutions

ICCV 2023poster

Convolutional neural network inference on video input is computationally expensive and requires high memory bandwidth. Recently, DeltaCNN managed to reduce the cost by only processing pixels with significant updates over the previous frame. However, DeltaCNN relies on static camera input. Moving cam…

Cited by 5PDFScholar
2023

Social Diffusion: Long-term Multiple Human Motion Anticipation

ICCV 2023poster

We propose Social Diffusion, a novel method for short-term and long-term forecasting of the motion of multiple persons as well as their social interactions. Jointly forecasting motions for multiple persons involved in social activities is inherently a challenging problem due to the interdependenci…

Cited by 20PDFcodeScholar
2022

DeltaCNN: End-to-End CNN Inference of Sparse Frame Differences in Videos

CVPR 2022poster

Convolutional neural network inference on video data requires powerful hardware for real-time processing. Given the inherent coherence across consecutive frames, large parts of a video typically change little. By skipping identical image regions and truncating insignificant pixel updates, computatio…

Cited by 34PDFcodeScholar
2022

Gradient Flow in Sparse Neural Networks and How Lottery Tickets Win

AAAI 2022technical

Sparse Neural Networks (NNs) can match the generalization of dense NNs using a fraction of the compute/storage for inference, and have the potential to enable efficient training. However, naively training unstructured sparse NNs from random initialization results in significantly worse generalizatio…

2022

Multiview Human Body Reconstruction from Uncalibrated Cameras

NeurIPS 2022accept

We present a new method to reconstruct 3D human body pose and shape by fusing visual features from multiview images captured by uncalibrated cameras. Existing multiview approaches often use spatial camera calibration (intrinsic and extrinsic parameters) to geometrically align and fuse visual feature…

Cited by 21SourcePDFScholar
2022

Neural Correspondence Field for Object Pose Estimation

ECCV 2022poster

"We propose a method for estimating the 6DoF pose of a rigid object with an available 3D model from a single RGB image. Unlike classical correspondence-based methods which predict 3D object coordinates at pixels of the input image, the proposed method predicts 3D object coordinates at 3D query point…

2021

HumanGPS: Geodesic PreServing Feature for Dense Human Correspondences

CVPR 2021poster

In this paper, we address the problem of building pixel-wise dense correspondences between human images under arbitrary camera viewpoints and body poses. Previous methods either assume small motions or rely on discriminative descriptors extracted from local patches, which cannot handle large motion…

Cited by 14PDFScholar
2021

Multiresolution Deep Implicit Functions for 3D Shape Representation

ICCV 2021poster

We introduce Multiresolution Deep Implicit Functions (MDIF), a hierarchical representation that can recover fine geometry detail, while being able to perform global operations such as shape completion. Our model represents a complex 3D shape with a hierarchy of latent grids, which can be decoded int…

Cited by 52PDFScholar
2020

Deep Implicit Volume Compression

CVPR 2020oral

We describe a novel approach for compressing truncated signed distance fields (TSDF) stored in 3D voxel grids, and their corresponding textures. To compress the TSDF, our method relies on a block-based neural network architecture trained end-to-end, achieving state-of-the-art rate-distortion trade-o…

Cited by 52PDFcodeScholar
2019

Volumetric Capture of Humans With a Single RGBD Camera via Semi-Parametric Learning

CVPR 2019poster

Volumetric (4D) performance capture is fundamental for AR/VR content generation. Whereas previous work in 4D performance capture has shown impressive results in studio settings, the technology is still far from being accessible to a typical consumer who, at best, might own a single RGBD sensor. Thus…

Cited by 47PDFScholar
2015

Learning an Efficient Model of Hand Shape Variation From Depth Images

CVPR 2015poster

We describe how to learn a compact and efficient model of the surface deformation of human hands. The model is built from a set of noisy depth images of a diverse set of subjects performing different poses with their hands. We represent the observed surface using Loop subdivision of a control mesh t…

Cited by 163SourcePDFScholar
2015

Opening the Black Box: Hierarchical Sampling Optimization for Estimating Human Hand Pose

ICCV 2015oral

We address the problem of hand pose estimation, formulated as an inverse problem. Typical approaches optimize an energy function over pose parameters using a `black box' image generation procedure. This procedure knows little about either the relationships between the parameters or the form of the…

Cited by 170PDFScholar