← Search

Jan-Michael Frahm

23 accepted papers

2024

Supervision Interpolation via LossMix: Generalizing Mixup for Object Detection and Beyond

AAAI 2024technical

The success of data mixing augmentations in image classification tasks has been well-received. However, these techniques cannot be readily applied to object detection due to challenges such as spatial misalignment, foreground/background distinction, and plurality of instances. To tackle these issues…

2023

A Practical Stereo Depth System for Smart Glasses

CVPR 2023poster

We present the design of a productionized end-to-end stereo depth sensing system that does pre-processing, online stereo rectification, and stereo depth estimation with a fallback to monocular depth estimation when rectification is unreliable. The output of our depth sensing system is then used in a…

Cited by 7SourcePDFScholar
2023

MVPSNet: Fast Generalizable Multi-view Photometric Stereo

ICCV 2023poster

We propose a fast and generalizable solution to Multiview Photometric Stereo (MVPS), called MVPSNet. The key to our approach is a feature extraction network that effectively combines images from the same view captured under multiple lighting conditions to extract geometric features from shading cues…

Cited by 17PDFScholar
2020

Boundary-Aware 3D Building Reconstruction From a Single Overhead Image

CVPR 2020poster

We propose a boundary-aware multi-task deep-learning-based framework for fast 3D building modeling from a single overhead image. Unlike most existing techniques which rely on multiple images for 3D scene modeling, we seek to model the buildings in the scene from a single overhead image by jointly le…

Cited by 61PDFcodeScholar
2020

VPLNet: Deep Single View Normal Estimation With Vanishing Points and Lines

CVPR 2020poster

We present a novel single-view surface normal estimation method that combines traditional line and vanishing point analysis with a deep learning approach. Starting from a color image and a Manhattan line map, we use a deep neural network to regress on a dense normal map, and a dense Manhattan label…

Cited by 47PDFScholar
2019

Recurrent Neural Network for (Un-)Supervised Learning of Monocular Video Visual Odometry and Depth

CVPR 2019poster

Deep learning-based, single-view depth estimation methods have recently shown highly promising results. However, such methods ignore one of the most important features for determining depth in the human vision system, which is motion. We propose a learning-based, multi-view dense depth map and odome…

Cited by 241PDFcodeScholar
2019

The Domain Transform Solver

CVPR 2019poster

We present a novel framework for edge-aware optimization that is an order of magnitude faster than the state of the art while maintaining comparable results. Our key insight is that the optimization can be formulated by leveraging properties of the domain transform, a method for edge-aware filtering…

Cited by 11PDFScholar
2018

Augmenting Crowd-Sourced 3D Reconstructions Using Semantic Detections

CVPR 2018poster

Image-based 3D reconstruction for Internet photo collections has become a robust technology to produce impressive virtual representations of real-world scenes. However, several fundamental challenges remain for Structure-from-Motion (SfM) pipelines, namely: the placement and reconstruction of transi…

Cited by 9SourcePDFScholar
2018

Rolling Shutter and Radial Distortion Are Features for High Frame Rate Multi-Camera Tracking

CVPR 2018poster

Traditionally, camera-based tracking approaches have treated rolling shutter and radial distortion as imaging artifacts that have to be overcome and corrected for in order to apply standard camera models and scene reconstruction methods. In this paper, we introduce a novel multi-camera tracking appr…

2016

From Dusk Till Dawn: Modeling in the Dark

CVPR 2016spotlight

Internet photo collections naturally contain a large variety of illumination conditions, with the largest difference between day and night images. Current modeling techniques do not embrace the broad illumination range often leading to reconstruction failure or severe artifacts. We present an algori…

Cited by 52PDFScholar
2015

From Single Image Query to Detailed 3D Reconstruction

CVPR 2015poster

Structure-from-Motion for unordered image collections has significantly advanced in scale over the last decade. This impressive progress can be in part attributed to the introduction of efficient retrieval methods for those systems. While this boosts scalability, it also limits the amount of detail…

Cited by 144SourcePDFScholar
2015

PAIGE: PAirwise Image Geometry Encoding for Improved Efficiency in Structure-From-Motion

CVPR 2015poster

Large-scale Structure-from-Motion systems typically spend major computational effort on pairwise image matching and geometric verification in order to discover connected components in large-scale, unordered image collections. In recent years, the research community has spent significant effort on im…

Cited by 44SourcePDFScholar
2015

Reconstructing the World* in Six Days *(As Captured by the Yahoo 100 Million Image Dataset)

CVPR 2015poster

We propose a novel, large-scale, structure-from-motion framework that advances the state of the art in data scalability from city-scale modeling (millions of images) to world-scale modeling (several tens of millions of images) using just a single computer. The main enabling technology is the use of…

Cited by 375SourcePDFScholar