← Search

Shahram Izadi

20 accepted papers

2020

Deep Implicit Volume Compression

CVPR 2020oral

We describe a novel approach for compressing truncated signed distance fields (TSDF) stored in 3D voxel grids, and their corresponding textures. To compress the TSDF, our method relies on a block-based neural network architecture trained end-to-end, achieving state-of-the-art rate-distortion trade-o…

Cited by 52PDFcodeScholar
2019

Volumetric Capture of Humans With a Single RGBD Camera via Semi-Parametric Learning

CVPR 2019poster

Volumetric (4D) performance capture is fundamental for AR/VR content generation. Whereas previous work in 4D performance capture has shown impressive results in studio settings, the technology is still far from being accessible to a typical consumer who, at best, might own a single RGBD sensor. Thus…

Cited by 47PDFScholar
2018

ActiveStereoNet: End-to-End Self-Supervised Learning for Active Stereo Systems

ECCV 2018poster

In this paper we present ActiveStereoNet, the first deep learning solution for active stereo systems. Due to the lack of ground truth, our method is fully self-supervised, yet it produces precise depth with a subpixel precision of 1/30th of a pixel; it does not suffer from the common over-smoothing…

Cited by 139SourcePDFScholar
2018

SOS: Stereo Matching in O(1) with Slanted Support Windows

IROS 2018poster

Depth cameras have accelerated research in many areas of computer vision. Most triangulation-based depth cameras, whether structured light systems like the Kinect or active (assisted) stereo systems, are based on the principle of stereo matching. Depth from stereo is an active research topic dating…

Cited by 24SourceScholar
2018

StereoNet: Guided Hierarchical Refinement for Real-Time Edge-Aware Depth Prediction

ECCV 2018poster

This paper presents StereoNet, the first end-to-end deep architecture for real-time stereo matching that runs at 60 fps on an NVidia Titan X, producing high-quality, edge-preserved, quantization-free depth maps. A key insight of this paper is that the network achieves a sub-pixel matching precision…

Cited by 461SourcePDFScholar
2017

DeepContext: Context-Encoding Neural Pathways for 3D Holistic Scene Understanding

ICCV 2017poster

3D context has been shown to be an extremely important cue for scene understanding, yet very little research has been done on integrating context information with deep models. This paper presents an approach to embed 3D context into the topology of a neural network trained to perform holistic scene…

Cited by 82PDFScholar
2017

Low Compute and Fully Parallel Computer Vision With HashMatch

ICCV 2017poster

Numerous computer vision problems such as stereo depth estimation, object-class segmentation and foreground/background segmentation can be formulated as per-pixel image labeling tasks. Given one or many images as input, the desired output of these methods is usually a spatially smooth assignment of…

Cited by 25PDFScholar
2017

UltraStereo: Efficient Learning-Based Matching for Active Stereo Systems

CVPR 2017spotlight

Efficient estimation of depth from pairs of stereo images is one of the core problems in computer vision. We efficiently solve the specialized problem of stereo matching under active illumination using a new learning-based algorithm. This type of 'active' stereo i.e. stereo matching where scene text…

Cited by 85PDFScholar
2016

Fits Like a Glove: Rapid and Reliable Hand Shape Personalization

CVPR 2016spotlight

We present a fast, practical method for personalizing a hand shape basis to an individual user's detailed hand shape using only a small set of depth images. To achieve this, we minimize an energy based on a sum of render-and-compare cost functions called the golden energy. However, this energy is on…

Cited by 154PDFScholar
2016

HyperDepth: Learning Depth From Structured Light Without Matching

CVPR 2016oral

Structured light sensors are popular due to their robustness to untextured scenes and multipath. These systems triangulate depth by solving a correspondence problem between each camera and projector pixel. This is often framed as a local stereo matching task, correlating patches of pixels in the obs…

Cited by 134PDFScholar
2015

3D Scanning Deformable Objects With a Single RGBD Sensor

CVPR 2015poster

We present a 3D scanning system for deformable objects that uses only a single Kinect sensor. Our work allows considerable amount of nonrigid deformations during scanning, and achieves high quality results without heavily constraining user or camera motion. We do not rely on any prior shape knowledg…

Cited by 178SourcePDFScholar
2015

A Light Transport Model for Mitigating Multipath Interference in Time-of-Flight Sensors

CVPR 2015poster

Continuous-wave Time-of-flight (TOF) range imaging has become a commercially viable technology with many applications in computer vision and graphics. However, the depth images obtained from TOF cameras contain scene dependent errors due to multipath interference (MPI). Specifically, MPI occurs when…

Cited by 105SourcePDFScholar
2015

Exploiting Uncertainty in Regression Forests for Accurate Camera Relocalization

CVPR 2015poster

Recent advances in camera relocalization use predictions from a regression forest to guide the camera pose optimization procedure. In these methods, each tree associates one pixel with a point in the scene's 3D world coordinate frame. In previous work, these predictions were point estimates and the…

Cited by 192SourcePDFScholar
2015

Incremental dense semantic stereo fusion for large-scale semantic scene reconstruction

ICRA 2015poster

Our abilities in scene understanding, which allow us to perceive the 3D structure of our surroundings and intuitively recognise the objects we see, are things that we largely take for granted, but for robots, the task of understanding large scenes quickly remains extremely challenging. Recently, sce…

Cited by 260SourceScholar
2015

Large-Scale and Drift-Free Surface Reconstruction Using Online Subvolume Registration

CVPR 2015poster

Depth cameras have helped commoditize 3D digitization of the real-world. It is now feasible to use a single Kinect-like camera to scan in an entire building or other large-scale scenes. At large scales, however, there is an inherent challenge of dealing with distortions and drift due to accumulated…

Cited by 80SourcePDFScholar
2015

Learning an Efficient Model of Hand Shape Variation From Depth Images

CVPR 2015poster

We describe how to learn a compact and efficient model of the surface deformation of human hands. The model is built from a set of noisy depth images of a diverse set of subjects performing different poses with their hands. We represent the observed surface using Loop subdivision of a control mesh t…

Cited by 163SourcePDFScholar