← Search

Robert Wang

9 accepted papers

2025

HOT3D: Hand and Object Tracking in 3D from Egocentric Multi-View Videos

CVPR 2025highlight

We introduce HOT3D, a publicly available dataset for egocentric hand and object tracking in 3D. The dataset offers over 833 minutes (3.7M+ images) of recordings that feature 19 subjects interacting with 33 diverse rigid objects. In addition to simple pick-up, observe, and put-down actions, the subje…

2024

emg2pose: A Large and Diverse Benchmark for Surface Electromyographic Hand Pose Estimation

NeurIPS 2024poster

Hands are the primary means through which humans interact with the world. Reliable and always-available hand pose inference could yield new and intuitive control schemes for human-computer interactions, particularly in virtual and augmented reality. Computer vision is effective but requires one or m…

2023

MotionDeltaCNN: Sparse CNN Inference of Frame Differences in Moving Camera Videos with Spherical Buffers and Padded Convolutions

ICCV 2023poster

Convolutional neural network inference on video input is computationally expensive and requires high memory bandwidth. Recently, DeltaCNN managed to reduce the cost by only processing pixels with significant updates over the previous frame. However, DeltaCNN relies on static camera input. Moving cam…

Cited by 5PDFScholar
2022

Assembly101: A Large-Scale Multi-View Video Dataset for Understanding Procedural Activities

CVPR 2022poster

Assembly101 is a new procedural activity dataset featuring 4321 videos of people assembling and disassembling 101 "take-apart" toy vehicles. Participants work without fixed instructions, and the sequences feature rich and natural variations in action ordering, mistakes, and corrections. Assembly101…

Cited by 246PDFcodeScholar
2022

DeltaCNN: End-to-End CNN Inference of Sparse Frame Differences in Videos

CVPR 2022poster

Convolutional neural network inference on video data requires powerful hardware for real-time processing. Given the inherent coherence across consecutive frames, large parts of a video typically change little. By skipping identical image regions and truncating insignificant pixel updates, computatio…

Cited by 34PDFcodeScholar
2022

Neural Correspondence Field for Object Pose Estimation

ECCV 2022poster

"We propose a method for estimating the 6DoF pose of a rigid object with an available 3D model from a single RGB image. Unlike classical correspondence-based methods which predict 3D object coordinates at pixels of the input image, the proposed method predicts 3D object coordinates at 3D query point…

2021

EM-POSE: 3D Human Pose Estimation From Sparse Electromagnetic Trackers

ICCV 2021poster

Fully immersive experiences in AR/VR depend on reconstructing the full body pose of the user without restricting their motion. In this paper we study the use of body-worn electromagnetic (EM) field-based sensing for the task of 3D human pose reconstruction. To this end, we present a method to estima…

Cited by 40PDFcodeScholar
2020

Lightweight Multi-View 3D Pose Estimation Through Camera-Disentangled Representation

CVPR 2020poster

We present a lightweight solution to recover 3D pose from multi-view images captured with spatially calibrated cameras. Building upon recent advances in interpretable representation learning, we exploit 3D geometry to fuse input images into a unified latent representation of pose, which is disentang…

Cited by 147PDFScholar