← Search

Daniel Cremers

181 accepted papers

2026

DynaTok: Token-Based 4D Reconstruction from Partial Point Clouds

ICML 2026poster

We address the problem of 4D reconstruction from partial point cloud sequences, where observations from depth sensors are incomplete, unordered, and lack explicit point correspondence over time. Recovering coherent 4D geometry in this geometry-only setting is challenging due to missing observations …

Cited by 0SourceScholar
2026

EgoFlow: Gradient-Guided Flow Matching for Egocentric 6DoF Object Motion Generation

CVPR 2026

Understanding and predicting object motion from egocentric video is fundamental to embodied perception and interaction. However, generating physically consistent 6DoF trajectories remains challenging due to occlusions, fast motion, and the lack of explicit physical reasoning in existing generative m

Cited by 0SourcecodeScholar
2026

From Pairwise Affinities to Functional Correspondences: Rethinking Attention

ICML 2026poster

Learning mappings between infinite-dimensional function spaces, or operator learning, is essential for many machine learning applications. Although transformer-based operators are popular, they often rely on token-wise attention. These methods treat continuous fields as discrete tokens and usually i…

Cited by 0SourceScholar
2026

GeodesicNVS: Probability Density Geodesic Flow Matching for Novel View Synthesis

CVPR 2026

Recent advances in generative modeling have substantially enhanced novel view synthesis, yet maintaining consistency across viewpoints remains challenging. Diffusion-based models rely on stochastic noise-to-data transitions, which obscure deterministic structures and yield inconsistent view predicti

Cited by 0SourceScholar
2026

Graph Neural Networks Are Not Continuous Across Graph Resolutions

ICML 2026poster

We show that contrary to conventional wisdom in the community, graph neural networks (GNNs) are not continuous with respect to all natural modes of graph convergence. As a result, GNNs may generate substantially different latent representations for graphs that are very similar. In particular they as…

Cited by 0SourceScholar
2026

HI-SLAM2: Geometry-Aware Gaussian SLAM for Fast Monocular Scene Reconstruction

ICRA 2026poster

We present HI-SLAM2, a geometry-aware Gaussian SLAM system that achieves fast and accurate monocular scene reconstruction using only RGB input. Existing Neural SLAM or 3DGS-based SLAM methods often trade off between rendering quality and geometry accuracy, our research demonstrates that both can be …

2026

HI-SLAM2: Geometry-Aware Gaussian SLAM for Fast Monocular Scene Reconstruction (Abstract Reprint)

AAAI 2026technical

We present HI-SLAM2, a geometry-aware Gaussian SLAM system that achieves fast and accurate monocular scene reconstruction using only RGB input. Existing Neural SLAM or 3DGS-based SLAM methods often trade off between rendering quality and geometry accuracy, our research demonstrates that both can be

Cited by 0SourcePDFScholar
2026

INSID3: Training-Free In-Context Segmentation with DINOv3

CVPR 2026

In-context segmentation (ICS) aims to segment arbitrary concepts, e.g., objects, parts, or personalized instances, given one annotated visual examples. Existing work relies on (i) fine-tuning vision foundation models (VFMs), which improves in-domain results but harms generalization, or (ii) combines

Cited by 0SourcecodeScholar
2026

Learning Eigenstructures of Unstructured Data Manifolds

CVPR 2026

We introduce a novel framework that directly learns a spectral basis for shape and manifold analysis from unstructured data, eliminating the need for traditional operator selection, discretization, and eigensolvers. Grounded in optimal-approximation theory, we train a network to decompose an implici

Cited by 0SourceScholar
2026

NOVA3R: Non-pixel-aligned Visual Transformer for Amodal 3D Reconstruction

ICLR 2026poster

We present NOVA3R, an effective approach for non-pixel-aligned 3D reconstruction from a set of unposed images, in a feed-forward manner. Unlike pixel-aligned methods that tie geometry to per-ray predictions, our formulation learns a global, view-agnostic scene representation that decouples reconstru…

Cited by 0SourcecodeScholar
2026

OccuFly: A 3D Vision Benchmark for Semantic Scene Completion from the Aerial Perspective

CVPR 2026

Semantic Scene Completion (SSC) is essential for 3D perception in mobile robotics, as it enables holistic scene understanding by jointly estimating dense volumetric occupancy and per-voxel semantics. Although SSC has been widely studied in terrestrial domains such as autonomous driving, aerial setti

Cited by 0SourcecodeScholar
2026

RINO: Rotation-Invariant Non-Rigid Correspondences

CVPR 2026

Dense 3D shape correspondence remains a central challenge in computer vision and graphics as many deep learning approaches still rely on intermediate geometric features or handcrafted descriptors, limiting their effectiveness under non-isometric deformations, partial data, and non-manifold inputs. T

Cited by 0SourceScholar
2026

Scene-Centric Unsupervised Video Panoptic Segmentation

CVPR 2026

Video panoptic segmentation (VPS) aims to jointly detect, segment, and track all objects while partitioning the video into semantically consistent regions. We introduce the task setting of unsupervised VPS, omitting any human supervision. Existing unsupervised scene understanding works mainly focuse

Cited by 0SourceScholar
2026

Teaching DINOv3 About Partial 3D Geometry: A Self-Supervised Geometry-Aware Approach

CVPR 2026

Partial shape matching is a crucial yet underexplored problem in 3D vision, with significant relevance to real-world scenarios where shapes are often only partially observed. Existing feature descriptors face difficulties in this setting, as traditional representations either struggle with the bound

Cited by 0SourcecodeScholar
2025

4Deform: Neural Surface Deformation for Robust Shape Interpolation

CVPR 2025poster

Generating realistic intermediate shapes between non-rigidly deformed shapes is a challenging task in computer vision, especially with unstructured data (e.g., point clouds) where temporal consistency across frames is lacking, and topologies are changing. Most interpolation methods are designed for…

Cited by 0SourcePDFScholar
2025

AnyCam: Learning to Recover Camera Poses and Intrinsics from Casual Videos

CVPR 2025poster

Estimating camera motion and intrinsics from casual videos is a core challenge in computer vision. Traditional bundle-adjustment based methods, such as SfM and SLAM, struggle to perform reliably on arbitrary data. Although specialized SfM approaches have been developed for handling dynamic scenes, t…

2025

Back on Track: Bundle Adjustment for Dynamic Scene Reconstruction

ICCV 2025poster

Traditional SLAM systems, which rely on bundle adjustment, struggle with the highly dynamic scenes commonly found in casual videos. Such videos entangle the motion of dynamic elements, undermining the assumption of static environments required by traditional systems. Existing techniques either filte…

Cited by 0SourcePDFScholar
2025

Dream-to-Recon: Monocular 3D Reconstruction with Diffusion-Depth Distillation from Single Images

ICCV 2025accepted

Volumetric scene reconstruction from a single image is crucial for a broad range of applications like autonomous driving and robotics. Recent volumetric reconstruction methods achieve impressive results, but generally require expensive 3D ground truth or multi-view supervision. We propose to leverag…

Cited by 0SourcePDFScholar
2025

EchoMatch: Partial-to-Partial Shape Matching via Correspondence Reflection

CVPR 2025poster

Finding correspondences between 3D shapes is a crucial problem in computer vision and graphics. While most research has focused on finding correspondences in settings where at least one of the shapes is complete, the realm of partial-to-partial shape matching remains under-explored. Yet, it is impor…

2025

Feed-Forward SceneDINO for Unsupervised Semantic Scene Completion

ICCV 2025poster

Semantic scene completion (SSC) aims to infer both the 3D geometry and semantics of a scene from single images. In contrast to prior work on SSC that heavily relies on expensive ground-truth annotations, we approach SSC in an unsupervised setting. Our novel method, SceneDINO, adapts techniques from…

2025

Finsler Multi-Dimensional Scaling: Manifold Learning for Asymmetric Dimensionality Reduction and Embedding

CVPR 2025poster

Dimensionality reduction is a fundamental task that aims to simplify complex data by reducing its feature dimensionality while preserving essential patterns, with core applications in data analysis and visualisation. To preserve the underlying data structure, multi-dimensional scaling (MDS) methods…

2025

GECO: Geometrically Consistent Embedding with Lightspeed Inference

ICCV 2025poster

Recent advancements in feature computation have revealed that self-supervised feature extractors can recognize semantic correspondences. However, these features often lack an understanding of objects' underlying 3D geometry. In this paper, we focus on learning features capable of semantically charac…

Cited by 0SourcePDFScholar
2025

Ground-Aware Automotive Radar Odometry

ICRA 2025

Odometry is crucial for the navigation of autonomous vehicles in unknown environments. While cameras and LiDARs are commonly used to estimate the ego-motion of a vehicle, these sensors face limitations under bad lighting and severe weather conditions. Automotive radars overcome these challenges, but

Cited by 5SourceScholar
2025

Higher-Order Ratio Cycles for Fast and Globally Optimal Shape Matching

CVPR 2025poster

In this work we address various shape matching problems that can be cast as finding cyclic paths in a product graph. This involves for example 2D-3D shape matching, 3D shape matching, or the matching of a contour to a graph. In this context, matchings are typically obtained as the minimum cost cycle…

2025

IPFormer: Visual 3D Panoptic Scene Completion with Context-Adaptive Instance Proposals

NeurIPS 2025poster

Semantic Scene Completion (SSC) has emerged as a pivotal approach for jointly learning scene geometry and semantics, enabling downstream applications such as navigation in mobile robotics. The recent generalization to Panoptic Scene Completion (PSC) advances the SSC domain by integrating instance-le…

Cited by 0SourceScholar
2025

Implicit Neural Surface Deformation with Explicit Velocity Fields

ICLR 2025poster

In this work, we introduce the first unsupervised method that simultaneously predicts time-varying neural implicit surfaces and deformations between pairs of point clouds. We propose to model the point movement using an explicit velocity field and directly deform a time-varying implicit field using…

2025

It's a (Blind) Match! Towards Vision-Language Correspondence without Parallel Data

CVPR 2025poster

The platonic representation hypothesis suggests that vision and language embeddings become more homogeneous as model and dataset sizes increase. In particular, pairwise distances within each modality become more similar. This suggests that as foundation models mature, it may become possible to match…

2025

Localizing Events in Videos with Multimodal Queries

CVPR 2025poster

Localizing events in videos based on semantic queries is a pivotal task in video understanding research and user-oriented applications like video search. Yet, current research predominantly relies on natural language queries (NLQs), overlooking the potential of using multimodal queries (MQs) that in…

Cited by 2SourcePDFScholar
2025

MA-DV${2}$F: A Multi-Agent Navigation Framework Using Dynamic Velocity Vector Field

RA-L 2025

In this paper, we propose MA-DV <inline-formula xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink"><tex-math notation="LaTeX">$^{2}$</tex-math></inline-formula> F: Multi-Agent Dynamic Velocity Vector Field. It is a framework for simultaneously controlling a gro

Cited by 1SourcecodeScholar
2025

MonoCT: Overcoming Monocular 3D Detection Domain Shift with Consistent Teacher Models

ICRA 2025

We tackle the problem of monocular 3D object detection across different sensors, environments, and camera setups. In this paper, we introduce a novel unsupervised domain adaptation approach, MonoCT, that generates highly accurate pseudo labels for self-supervision. Inspired by our observation that a

Cited by 9SourceScholar
2025

Nonisotropic Gaussian Diffusion for Realistic 3D Human Motion Prediction

CVPR 2025poster

Probabilistic human motion prediction aims to forecast multiple possible future movements from past observations. While current approaches report high diversity and realism, they often generate motions with undetected limb stretching and jitter. To address this, we introduce SkeletonDiffusion, a lat…

2025

OPAL: Visibility-aware LiDAR-to-OpenStreetMap Place Recognition via Adaptive Radial Fusion

CoRL 2025poster

LiDAR place recognition is a critical capability for autonomous navigation and cross-modal localization in large-scale outdoor environments. Existing approaches predominantly depend on pre-built 3D dense maps or aerial imagery, which impose significant storage overhead and lack real-time adaptabilit…

Cited by 0SourceScholar
2025

OrthoLoC: UAV 6-DoF Localization and Calibration Using Orthographic Geodata

NeurIPS 2025oral

Accurate visual localization from aerial views is a fundamental problem with applications in mapping, large-area inspection, and search-and-rescue operations. In many scenarios, these systems require high-precision localization while operating with limited resources (e.g., no internet connection or…

Cited by 0SourcecodeScholar
2025

Scene-Centric Unsupervised Panoptic Segmentation

CVPR 2025highlight

Unsupervised panoptic segmentation aims to partition an image into semantically meaningful regions and distinct object instances without training on manually annotated data. In contrast to prior work on unsupervised panoptic scene understanding, we eliminate the need for object-centric training data…

2025

Semantic Library Adaptation: LoRA Retrieval and Fusion for Open-Vocabulary Semantic Segmentation

CVPR 2025poster

Open-vocabulary semantic segmentation models associate vision and text to label pixels from an undefined set of classes using textual queries, providing versatile performance on novel datasets. However, large shifts between training and test domains degrade their performance, requiring fine-tuning f…

2025

SparseAlign: a Fully Sparse Framework for Cooperative Object Detection

CVPR 2025poster

Cooperative perception can increase the view field and decrease the occlusion of an ego vehicle, hence improving the perception performance and safety of autonomous driving. Despite the success of previous works on cooperative object detection, they mostly operate on dense Bird's Eye View (BEV) feat…

Cited by 0SourcePDFScholar
2025

TrafficLoc: Localizing Traffic Surveillance Cameras in 3D Scenes

ICCV 2025poster

We tackle the problem of localizing traffic cameras within a 3D reference map and propose a novel image-to-point cloud registration (I2P) method, TrafficLoc, in a coarse-to-fine matching fashion. To overcome the lack of large-scale real-world intersection datasets, we first introduce Carla Intersect…

2025

TurnBack: A Geospatial Route Cognition Benchmark for Large Language Models through Reverse Route

EMNLP 2025

Humans can interpret geospatial information through natural language, while the geospatial cognition capabilities of Large Language Models (LLMs) remain underexplored. Prior research in this domain has been constrained by non-quantifiable metrics, limited evaluation datasets; unclear research hierar

Cited by 0SourcePDFScholar
2025

UrbanIng-V2X: A Large-Scale Multi-Vehicle, Multi-Infrastructure Dataset Across Multiple Intersections for Cooperative Perception

NeurIPS 2025poster

Recent cooperative perception datasets have played a crucial role in advancing smart mobility applications by enabling information exchange between intelligent agents, helping to overcome challenges such as occlusions and improving overall scene understanding. While some existing real-world datasets…

Cited by 0SourcecodeScholar
2025

VoxNeRF: Bridging Voxel Representation and Neural Radiance Fields for Enhanced Indoor View Synthesis

RA-L 2025

The generation of high-fidelity view synthesis is essential for robotic navigation and interaction but remains challenging, particularly in indoor environments and real-time scenarios. Existing techniques often require significant computational resources for both training and rendering, and they fre

Cited by 2SourceScholar
2024

An Analytical Solution to Gauss-Newton Loss for Direct Image Alignment

ICLR 2024oral

Direct image alignment is a widely used technique for relative 6DoF pose estimation between two images, but its accuracy strongly depends on pose initialization. Therefore, recent end-to-end frameworks increase the convergence basin of the learned feature descriptors with special training objectives…

Cited by 0SourcePDFScholar
2024

An Image is Worth 32 Tokens for Reconstruction and Generation

NeurIPS 2024poster

Recent advancements in generative models have highlighted the crucial role of image tokenization in the efficient synthesis of high-resolution images. Tokenization, which transforms images into latent representations, reduces computational demands compared to directly processing pixels and enhances…

2024

Boosting Self-Supervision for Single-View Scene Completion via Knowledge Distillation

CVPR 2024poster

Inferring scene geometry from images via Structure from Motion is a long-standing and fundamental problem in computer vision. While classical approaches and more recently depth map predictions only focus on the visible parts of a scene the task of scene completion aims to reason about geometry even…

Cited by 2SourcePDFScholar
2024

Cache Me if You Can: Accelerating Diffusion Models through Block Caching

CVPR 2024poster

Diffusion models have recently revolutionized the field of image synthesis due to their ability to generate photorealistic images. However one of the major drawbacks of diffusion models is that the image generation process is costly. A large image-to-image network has to be applied many times to ite…

Cited by 51SourcePDFScholar
2024

Enhancing Hypergradients Estimation: A Study of Preconditioning and Reparameterization

AISTATS 2024poster

Bilevel optimization aims to optimize an outer objective function that depends on the solution to an inner optimization problem. It is routinely used in Machine Learning, notably for hyperparameter tuning. The conventional method to compute the so-called hypergradient of the outer problem is to use…

2024

Finsler-Laplace-Beltrami Operators with Application to Shape Analysis

CVPR 2024poster

The Laplace-Beltrami operator (LBO) emerges from studying manifolds equipped with a Riemannian metric. It is often called the swiss army knife of geometry processing as it allows to capture intrinsic shape information and gives rise to heat diffusion geodesic distances and a multitude of shape descr…

Cited by 8SourcePDFScholar
2024

Flattening the Parent Bias: Hierarchical Semantic Segmentation in the Poincare Ball

CVPR 2024poster

Hierarchy is a natural representation of semantic taxonomies including the ones routinely used in image segmentation. Indeed recent work on semantic segmentation reports improved accuracy from supervised training leveraging hierarchical label structures. Encouraged by these results we revisit the fu…

2024

GlobalPointer: Large-Scale Plane Adjustment with Bi-Convex Relaxation

ECCV 2024poster

"Plane adjustment (PA) is crucial for many 3D applications, involving simultaneous pose estimation and plane recovery. Despite recent advancements, it remains a challenging problem in the realm of multi-view point cloud registration. Current state-of-the-art methods can achieve globally optimal conv…

2024

MICDrop: Masking Image and Depth Features via Complementary Dropout for Domain-Adaptive Semantic Segmentation

ECCV 2024poster

"Unsupervised Domain Adaptation (UDA) is the task of bridging the domain gap between a labeled source domain, e.g., synthetic data, and an unlabeled target domain. We observe that current UDA methods show inferior results on fine structures and tend to oversegment objects with ambiguous appearance.…

2024

Partial-to-Partial Shape Matching with Geometric Consistency

CVPR 2024poster

Finding correspondences between 3D shapes is an important and long-standing problem in computer vision graphics and beyond. A prominent challenge are partial-to-partial shape matching settings which occur when the shapes to match are only observed incompletely (e.g. from 3D scanning). Although parti…

2024

Physically-Based Photometric Bundle Adjustment in Non-Lambertian Environments

IROS 2024poster

Photometric bundle adjustment (PBA) is widely used in estimating the camera pose and 3D geometry by assuming a Lambertian world. However, the assumption of photometric consistency is often violated since the non-diffuse reflection is common in real-world environments. The photometric inconsistency s…

Cited by 0SourceScholar
2024

Power Variable Projection for Initialization-Free Large-Scale Bundle Adjustment

ECCV 2024poster

"Most Bundle Adjustment (BA) solvers like the Levenberg-Marquard algorithm require a good initialization. Instead, initialization-free BA remains a largely uncharted territory. The under-explored Variable Projection algorithm (VarPro) exhibits a wide convergence basin even without initialization. Co…

2024

Quality-Aware Translation Models: Efficient Generation and Quality Estimation in a Single Model

ACL 2024long

Maximum-a-posteriori (MAP) decoding is the most widely used decoding strategy for neural machine translation (NMT) models. The underlying assumption is that model probability correlates well with human judgment, with better translations getting assigned a higher score by the model. However, research…

2024

SatSynth: Augmenting Image-Mask Pairs through Diffusion Models for Aerial Semantic Segmentation

CVPR 2024poster

In recent years semantic segmentation has become a pivotal tool in processing and interpreting satellite imagery. Yet a prevalent limitation of supervised learning techniques remains the need for extensive manual annotations by experts. In this work we explore the potential of generative image diffu…

Cited by 23SourcePDFScholar
2024

Sparse Views Near Light: A Practical Paradigm for Uncalibrated Point-light Photometric Stereo

CVPR 2024poster

Neural approaches have shown a significant progress on camera-based reconstruction. But they require either a fairly dense sampling of the viewing sphere or pre-training on an existing dataset thereby limiting their generalizability. In contrast photometric stereo (PS) approaches have shown great po…

Cited by 6SourcePDFScholar
2024

Spectral Meets Spatial: Harmonising 3D Shape Matching and Interpolation

CVPR 2024poster

Although 3D shape matching and interpolation are highly interrelated they are often studied separately and applied sequentially to relate different 3D shapes thus resulting in sub-optimal performance. In this work we present a unified framework to predict both point-wise correspondences and shape in…

Cited by 10SourcePDFScholar
2024

Text2Loc: 3D Point Cloud Localization from Natural Language

CVPR 2024poster

We tackle the problem of 3D point cloud localization based on a few natural linguistic descriptions and introduce a novel neural network Text2Loc that fully interprets the semantic relationship between points and text. Text2Loc follows a coarse-to-fine localization pipeline: text-submap global place…

2024

Variational Learning is Effective for Large Deep Networks

ICML 2024spotlight

We give extensive empirical evidence against the common belief that variational learning is ineffective for large neural networks. We show that an optimizer called Improved Variational Online Newton (IVON) consistently matches or outperforms Adam for training large networks such as GPT-2 and ResNets…

2023

Behind the Scenes: Density Fields for Single View Reconstruction

CVPR 2023poster

Inferring a meaningful geometric scene representation from a single image is a fundamental problem in computer vision. Approaches based on traditional depth map prediction can only reason about areas that are visible in the image. Currently, neural radiance fields (NeRFs) can capture true 3D includi…

2023

Beyond In-Domain Scenarios: Robust Density-Aware Calibration

ICML 2023poster

Calibrating deep learning models to yield uncertainty-aware predictions is crucial as deep neural networks get increasingly deployed in safety-critical applications. While existing post-hoc calibration methods achieve impressive results on in-domain test datasets, they are limited by their inability…

2023

CASSPR: Cross Attention Single Scan Place Recognition

ICCV 2023poster

Place recognition based on point clouds (LiDAR) is an important component for autonomous robots or self-driving vehicles. Current SOTA performance is achieved on accumulated LiDAR submaps using either point-based or voxel-based structures. While voxel-based approaches nicely integrate spatial contex…

Cited by 68PDFcodeScholar
2023

DDIT: Semantic Scene Completion via Deformable Deep Implicit Templates

ICCV 2023poster

Scene reconstructions are often incomplete due to occlusions and limited viewpoints. There have been efforts to use semantic information for scene completion. However, the completed shapes may be rough and imprecise since respective methods rely on 3D convolution and/or lack effective shape constrai…

Cited by 10PDFScholar
2023

E-NeRF: Neural Radiance Fields From a Moving Event Camera

RA-L 2023

Estimating neural radiance fields (NeRFs) from “ideal” images has been extensively studied in the computer vision community. Most approaches assume optimal illumination and slow camera motion. These assumptions are often violated in robotic applications, where images may contain motion blur, and the

Cited by 102SourcecodeScholar
2023

G-MSM: Unsupervised Multi-Shape Matching With Graph-Based Affinity Priors

CVPR 2023poster

We present G-MSM (Graph-based Multi-Shape Matching), a novel unsupervised learning approach for non-rigid shape correspondence. Rather than treating a collection of input poses as an unordered set of samples, we explicitly model the underlying shape data manifold. To this end, we propose an adaptive…

2023

Learning Correspondence Uncertainty via Differentiable Nonlinear Least Squares

CVPR 2023poster

We propose a differentiable nonlinear least squares framework to account for uncertainty in relative pose estimation from feature correspondences. Specifically, we introduce a symmetric version of the probabilistic normal epipolar constraint, and an approach to estimate the covariance of feature pos…

Cited by 11SourcePDFScholar
2023

Learning Expressive Priors for Generalization and Uncertainty Estimation in Neural Networks

ICML 2023poster

In this work, we propose a novel prior learning method for advancing generalization and uncertainty estimation in deep neural networks. The key idea is to exploit scalable and structured posteriors of neural networks as informative priors with generalization guarantees. Our learned priors provide ex…

2023

Power Bundle Adjustment for Large-Scale 3D Reconstruction

CVPR 2023poster

We introduce Power Bundle Adjustment as an expansion type algorithm for solving large-scale bundle adjustment problems. It is based on the power series expansion of the inverse Schur complement and constitutes a new family of solvers that we call inverse expansion methods. We theoretically justify t…

2023

Robust Autonomous Vehicle Pursuit Without Expert Steering Labels

RA-L 2023

In this work, we present a learning method for both lateral and longitudinal motion control of an ego-vehicle for the task of vehicle pursuit. The car being controlled does not have a pre-defined route, rather it reactively adapts to follow a target vehicle while maintaining a safety distance. To tr

Cited by 1SourceScholar
2023

SIGMA: Scale-Invariant Global Sparse Shape Matching

ICCV 2023poster

We propose a novel mixed-integer programming (MIP) formulation for generating precise sparse correspondences for highly non-rigid shapes. To this end, we introduce a projected Laplace-Beltrami operator (PLBO) which combines intrinsic and extrinsic geometric information to measure the deformation qua…

Cited by 10PDFScholar
2023

Semidefinite Relaxations for Robust Multiview Triangulation

CVPR 2023poster

We propose an approach based on convex relaxations for certifiably optimal robust multiview triangulation. To this end, we extend existing relaxation approaches to non-robust multiview triangulation by incorporating a least squares cost function. We propose two formulations, one based on epipolar co…

2023

To Adapt or Not to Adapt? Real-Time Adaptation for Semantic Segmentation

ICCV 2023poster

The goal of Online Domain Adaptation for semantic segmentation is to handle unforeseeable domain changes that occur during deployment, like sudden weather events. However, the high computational costs associated with brute-force adaptation make this paradigm unfeasible for real-world applications. I…

Cited by 13PDFcodeScholar
2022

A Scalable Combinatorial Solver for Elastic Geometrically Consistent 3D Shape Matching

CVPR 2022poster

We present a scalable combinatorial algorithm for globally optimizing over the space of geometrically consistent mappings between 3D shapes. We use the mathematically elegant formalism proposed by Windheuser et al. (ICCV, 2011) where 3D shape matching was formulated as an integer linear program over…

Cited by 22PDFcodeScholar
2022

A Unified Framework for Implicit Sinkhorn Differentiation

CVPR 2022poster

The Sinkhorn operator has recently experienced a surge of popularity in computer vision and related fields. One major reason is its ease of integration into deep learning frameworks. To allow for an efficient training of respective neural networks, we propose an algorithm that obtains analytical gra…

Cited by 26PDFcodeScholar
2022

DirectTracker: 3D Multi-Object Tracking Using Direct Image Alignment and Photometric Bundle Adjustment

IROS 2022poster

Direct methods have shown excellent performance in the applications of visual odometry and SLAM. In this work we propose to leverage their effectiveness for the task of 3D multi-object tracking. To this end, we propose DirectTracker, a framework that effectively combines direct image alignment for t…

Cited by 5SourceScholar
2022

DynamicEarthNet: Daily Multi-Spectral Satellite Dataset for Semantic Change Segmentation

CVPR 2022poster

Earth observation is a fundamental tool for monitoring the evolution of land use in specific areas of interest. Observing and precisely defining change, in this context, requires both time-series data and pixel-wise segmentations. To that end, we propose the DynamicEarthNet dataset that consists of…

Cited by 107PDFScholar
2022

Gradient-SDF: A Semi-Implicit Surface Representation for 3D Reconstruction

CVPR 2022poster

We present Gradient-SDF, a novel representation for 3D geometry that combines the advantages of implict and explicit representations. By storing at every voxel both the signed distance field as well as its gradient vector field, we enhance the capability of implicit representations with approaches o…

Cited by 21PDFScholar
2022

Intrinsic Neural Fields: Learning Functions on Manifolds

ECCV 2022poster

"Neural fields have gained significant attention in the computer vision community due to their excellent performance in novel view synthesis, geometry reconstruction, and generative modeling. Some of their advantages are a sound theoretic foundation and an easy implementation in current deep learnin…

2022

Joint Deep Multi-Graph Matching and 3D Geometry Learning from Inhomogeneous 2D Image Collections

AAAI 2022technical

Graph matching aims to establish correspondences between vertices of graphs such that both the node and edge attributes agree. Various learning-based methods were recently proposed for finding correspondences between image key points based on deep graph matching formulations. While these approaches…

Cited by 8SourcePDFScholar
2022

PRISM: Probabilistic Real-Time Inference in Spatial World Models

CoRL 2022oral

We introduce PRISM, a method for real-time filtering in a probabilistic generative model of agent motion and visual perception. Previous approaches either lack uncertainty estimates for the map and agent state, do not run in real-time, do not have a dense scene representation or do not model agent d…

Cited by 2SourceScholar
2022

Parameterized Temperature Scaling for Boosting the Expressive Power in Post-Hoc Uncertainty Calibration

ECCV 2022poster

"We address the problem of uncertainty calibration and introduce a novel calibration method, Parametrized Temperature Scaling (PTS). Standard deep neural networks typically yield uncalibrated predictions, which can be transformed into calibrated confidence scores using post-hoc calibration methods.…

2022

The Probabilistic Normal Epipolar Constraint for Frame-to-Frame Rotation Optimization Under Uncertain Feature Positions

CVPR 2022poster

The estimation of the relative pose of two camera views is a fundamental problem in computer vision. Kneip et al. proposed to solve this problem by introducing the normal epipolar constraint (NEC). However, their approach does not take into account uncertainties, so that the accuracy of the estimate…

Cited by 11PDFScholar
2022

Vision-Based Large-scale 3D Semantic Mapping for Autonomous Driving Applications

ICRA 2022poster

In this paper, we present a complete pipeline for 3D semantic mapping solely based on a stereo camera system. The pipeline comprises a direct sparse visual odometry frontend as well as a back-end for global optimization including GNSS integration, and semantic 3D point cloud labeling. We propose a s…

Cited by 10SourceScholar
2022

What Makes Graph Neural Networks Miscalibrated?

NeurIPS 2022accept

Given the importance of getting calibrated predictions and reliable uncertainty estimations, various post-hoc calibration methods have been developed for neural networks on standard multi-class classification tasks. However, these methods are not well suited for calibrating graph neural networks (GN…

2021

Explicit pairwise factorized graph neural network for semi-supervised node classification

UAI 2021poster

Node features and structural information of a graph are both crucial for semi-supervised node classification problems. A variety of graph neural network (GNN) based approaches have been proposed to tackle these problems, which typically determine output labels through feature aggregation. This can b…

2021

MonoRec: Semi-Supervised Dense Reconstruction in Dynamic Environments From a Single Moving Camera

CVPR 2021poster

In this paper, we propose MonoRec, a semi-supervised monocular dense reconstruction architecture that predicts depth maps from a single moving camera in dynamic environments. MonoRec is based on a multi-view stereo setting which encodes the information of multiple consecutive images in a cost volume…

Cited by 105PDFcodeScholar
2021

NeuroMorph: Unsupervised Shape Interpolation and Correspondence in One Go

CVPR 2021poster

We present NeuroMorph, a new neural network architecture that takes as input two 3D shapes and produces in one go, i.e. in a single feed forward pass, a smooth interpolation and point-to-point correspondences between them. The interpolation, expressed as a deformation field, changes the pose of the…

Cited by 81PDFScholar
2021

Post-Hoc Uncertainty Calibration for Domain Drift Scenarios

CVPR 2021poster

We address the problem of uncertainty calibration. While standard deep neural networks typically yield uncalibrated predictions, calibrated confidence scores that are representative of the true likelihood of a prediction can be achieved using post-hoc calibration methods. However, to date, the focus…

Cited by 89PDFcodeScholar
2021

SOE-Net: A Self-Attention and Orientation Encoding Network for Point Cloud Based Place Recognition

CVPR 2021poster

We tackle the problem of place recognition from point cloud data and introduce a self-attention and orientation encoding network (SOE-Net) that fully explores the relationship between points and incorporates long-range context into point-wise local descriptors. Local information of each point from e…

Cited by 189PDFcodeScholar
2021

STEP: Segmenting and Tracking Every Pixel

NeurIPS 2021poster

The task of assigning semantic classes and track identities to every pixel in a video is called video panoptic segmentation. Our work is the first that targets this task in a real-world setting requiring dense interpretation in both spatial and temporal domains. As the ground-truth for this task is…

Cited by 89SourcecodeScholar
2021

Self-Supervised Steering Angle Prediction for Vehicle Control Using Visual Odometry

AISTATS 2021poster

Vision-based learning methods for self-driving cars have primarily used supervised approaches that require a large number of labels for training. However, those labels are usually difficult and expensive to obtain. In this paper, we demonstrate how a model can be trained to control a vehicle’s traje…

Cited by 3SourcePDFScholar
2021

Sparse Quadratic Optimisation over the Stiefel Manifold with Application to Permutation Synchronisation

NeurIPS 2021poster

We address the non-convex optimisation problem of finding a sparse matrix on the Stiefel manifold (matrices with mutually orthogonal columns of unit length) that maximises (or minimises) a quadratic objective function. Optimisation problems on the Stiefel manifold occur for example in spectral relax…

Cited by 12SourcePDFScholar
2021

Square Root Bundle Adjustment for Large-Scale Reconstruction

CVPR 2021poster

We propose a new formulation for the bundle adjustment problem which relies on nullspace marginalization of landmark variables by QR decomposition. Our approach, which we call square root bundle adjustment, is algebraically equivalent to the commonly used Schur complement trick, improves the numeric…

Cited by 27PDFScholar
2021

Square Root Marginalization for Sliding-Window Bundle Adjustment

ICCV 2021poster

In this paper we propose a novel square root sliding-window bundle adjustment suitable for real-time odometry applications. The square root formulation pervades three major aspects of our optimization-based sliding-window estimator: for bundle adjustment we eliminate landmark variables with nullspac…

Cited by 19PDFScholar
2021

TANDEM: Tracking and Dense Mapping in Real-time using Deep Multi-view Stereo

CoRL 2021poster

In this paper, we present TANDEM a real-time monocular tracking and dense mapping framework. For pose estimation, TANDEM performs photometric bundle adjustment based on a sliding window of keyframes. To increase the robustness, we propose a novel tracking front-end that performs dense direct image a…

Cited by 91SourcecodeScholar
2021

Tight Integration of Feature-based Relocalization in Monocular Direct Visual Odometry

ICRA 2021poster

In this paper we propose a framework for inte-grating map-based relocalization into online direct visual odometry. To achieve map-based relocalization for direct methods, we integrate image features into Direct Sparse Odometry (DSO) and rely on feature matching to associate online visual odometry (V…

Cited by 14SourceScholar
2021

Variational Data Assimilation with a Learned Inverse Observation Operator

ICML 2021spotlight

Variational data assimilation optimizes for an initial state of a dynamical system such that its evolution fits observational data. The physical model can subsequently be evolved into the future to make predictions. This principle is a cornerstone of large scale forecasting applications such as nume…

2021

Vision-Based Mobile Robotics Obstacle Avoidance With Deep Reinforcement Learning

ICRA 2021poster

Obstacle avoidance is a fundamental and challenging problem for autonomous navigation of mobile robots. In this paper, we consider the problem of obstacle avoidance in simple 3D environments where the robot has to solely rely on a single monocular camera. In particular, we are interested in solving…

Cited by 58SourceScholar
2021

i3DMM: Deep Implicit 3D Morphable Model of Human Heads

CVPR 2021poster

We present the first deep implicit 3D morphable model (i3DMM) of full heads. Unlike earlier morphable face models it not only captures identity-specific geometry, texture, and expressions of the frontal face, but also models the entire head, including hair. We collect a new dataset consisting of 64…

Cited by 139PDFcodeScholar
2020

Correspondence-Free Material Reconstruction using Sparse Surface Constraints

CVPR 2020poster

We present a method to infer physical material parameters, and even external boundaries, from the scanned motion of a homogeneous deformable object via the solution of an inverse problem. Parameters are estimated from real-world data sources such as sparse observations from a Kinect sensor without c…

Cited by 18PDFScholar
2020

D3VO: Deep Depth, Deep Pose and Deep Uncertainty for Monocular Visual Odometry

CVPR 2020oral

We propose D3VO as a novel framework for monocular visual odometry that exploits deep networks on three levels -- deep depth, pose and uncertainty estimation. We first propose a novel self-supervised monocular depth estimation network trained on stereo videos without any external supervision. In par…

Cited by 519PDFScholar
2020

DH3D: Deep Hierarchical 3D Descriptors for Robust Large-Scale 6DoF Relocalization

ECCV 2020poster

For relocalization in large-scale point clouds, we propose the first approach that unifies global place recognition and local 6DoF pose refinement. To this end, we design a Siamese network that jointly learns 3D local feature detection and description directly from raw 3D points. It integrates FlexC…

2020

Deep Shells: Unsupervised Shape Correspondence with Optimal Transport

NeurIPS 2020poster

We propose a novel unsupervised learning approach to 3D shape correspondence that builds a multiscale matching pipeline into a deep neural network. This approach is based on smooth shells, the current state-of-the-art axiomatic correspondence method, which requires an a priori stochastic search over…

2020

DirectShape: Direct Photometric Alignment of Shape Priors for Visual Vehicle Pose and Shape Estimation

ICRA 2020poster

Scene understanding from images is a challenging problem encountered in autonomous driving. On the object level, while 2D methods have gradually evolved from computing simple bounding boxes to delivering finer grained results like instance segmentations, the 3D family is still dominated by estimatin…

Cited by 40SourceScholar
2020

Efficient Derivative Computation for Cumulative B-Splines on Lie Groups

CVPR 2020oral

Continuous-time trajectory representation has recently gained popularity for tasks where the fusion of high-frame-rate sensors and multiple unsynchronized devices is required. Lie group cumulative B-splines are a popular way of representing continuous trajectories without singularities. They have be…

Cited by 103PDFcodeScholar
2020

From Planes to Corners: Multi-Purpose Primitive Detection in Unorganized 3D Point Clouds

RA-L 2020

We propose anew method for segmentation-free joint estimation of orthogonal planes, their intersection lines, relationship graph and corners lying at the intersection of three orthogonal planes. Such unified scene exploration under orthogonality allows for multitudes of applications such as semantic

Cited by 12SourcecodeScholar
2020

Optimization of Graph Total Variation via Active-Set-based Combinatorial Reconditioning

AISTATS 2020poster

Structured convex optimization on weighted graphs finds numerous applications in machine learning and computer vision. In this work, we propose a novel adaptive preconditioning strategy for proximal algorithms on this problem class. Our preconditioner is driven by a sharp analysis of the local linea…

Cited by 4SourcePDFScholar
2020

PrimiTect: Fast Continuous Hough Voting for Primitive Detection

ICRA 2020poster

This paper tackles the problem of data abstraction in the context of 3D point sets. Our method classifies points into different geometric primitives, such as planes and cones, leading to a compact representation of the data. Being based on a semi-global Hough voting scheme, the method does not need…

Cited by 19SourceScholar
2020

Visual-Inertial Mapping With Non-Linear Factor Recovery

RA-L 2020

Cameras and inertial measurement units are complementary sensors for ego-motion estimation and environment mapping. Their combination makes visual-inertial odometry (VIO) systems more accurate and robust. For globally consistent mapping, however, combining visual and inertial information is not stra

Cited by 228SourceScholar
2019

Lifting Vectorial Variational Problems: A Natural Formulation Based on Geometric Measure Theory and Discrete Exterior Calculus

CVPR 2019oral

Numerous tasks in imaging and vision can be formulated as variational problems over vector-valued maps. We approach the relaxation and convexification of such vectorial variational problems via a lifting to the space of currents. To that end, we recall that functionals with polyconvex Lagrangians c…

Cited by 16PDFScholar
2019

Optimization of Inf-Convolution Regularized Nonconvex Composite Problems

AISTATS 2019poster

In this work, we consider nonconvex composite problems that involve inf-convolution with a Legendre function, which gives rise to an anisotropic generalization of the proximal mapping and Moreau-envelope. In a convex setting such problems can be solved via alternating minimization of a splitting for…

Cited by 7SourcePDFScholar
2019

Rolling-Shutter Modelling for Direct Visual-Inertial Odometry

IROS 2019poster

We present a direct visual-inertial odometry (VIO) method which estimates the motion of the sensor setup and sparse 3D geometry of the environment based on measurements from a rolling-shutter camera and an inertial measurement unit (IMU). The visual part of the system performs a photometric bundle a…

Cited by 45SourceScholar
2019

Towards Generalizing Sensorimotor Control Across Weather Conditions

IROS 2019poster

The ability of deep learning models to generalize well across different scenarios depends primarily on the quality and quantity of annotated data. Labeling large amounts of data for all possible scenarios that a model may encounter would not be feasible; if even possible. We propose a framework to d…

Cited by 7SourceScholar
2019

Variational Uncalibrated Photometric Stereo Under General Lighting

ICCV 2019poster

Photometric stereo (PS) techniques nowadays remain constrained to an ideal laboratory setup where modeling and calibration of lighting is amenable. To eliminate such restrictions, we propose an efficient principled variational approach to uncalibrated PS under general illumination. To this end, the…

Cited by 44PDFcodeScholar
2018

A Nonconvex Proximal Splitting Algorithm under Moreau-Yosida Regularization

AISTATS 2018poster

We tackle highly nonconvex, nonsmooth composite optimization problems whose objectives comprise a Moreau-Yosida regularized term. Classical nonconvex proximal splitting algorithms, such as nonconvex ADMM, suffer from lack of convergence for such a problem class. To overcome this difficulty, in this…

Cited by 0SourcePDFScholar
2018

Challenges in Monocular Visual Odometry: Photometric Calibration, Motion Bias, and Rolling Shutter Effect

RA-L 2018

Monocular visual odometry (VO) and simultaneous localization and mapping (SLAM) have seen tremendous improvements in accuracy, robustness, and efficiency, and have gained increasing popularity over recent years. Nevertheless, not so many discussions have been carried out to reveal the influences of

Cited by 117SourceScholar
2018

Combinatorial Preconditioners for Proximal Algorithms on Graphs

AISTATS 2018poster

We present a novel preconditioning technique for proximal optimization methods that relies on graph algorithms to construct effective preconditioners. Such combinatorial preconditioners arise from partitioning the graph into forests. We prove that certain decompositions lead to a theoretically optim…

Cited by 0SourcePDFScholar
2018

Deep Virtual Stereo Odometry: Leveraging Deep Depth Prediction for Monocular Direct Sparse Odometry

ECCV 2018poster

Monocular visual odometry approaches that purely rely on geometric cues are prone to scale drift and require sufficient motion parallax in successive frames for motion estimation and 3D reconstruction. In this paper, we propose to leverage deep monocular depth prediction to overcome limitations of g…

Cited by 427SourcePDFScholar
2018

Direct Sparse Odometry With Rolling Shutter

ECCV 2018poster

Neglecting the effects of rolling-shutter cameras for visual odometry (VO) severely degrades accuracy and robustness. In this paper, we propose a novel direct monocular VO method that incorporates a rolling-shutter model. Our approach extends direct sparse odometry which performs direct bundle adjus…

Cited by 56SourcePDFScholar
2018

Direct Sparse Visual-Inertial Odometry Using Dynamic Marginalization

ICRA 2018poster

We present VI-DSO, a novel approach for visual-inertial odometry, which jointly estimates camera poses and sparse scene geometry by minimizing photometric and IMU measurement errors in a combined energy functional. The visual part of the system performs a bundle-adjustment like optimization on a spa…

Cited by 322SourcecodeScholar
2018

Discrete-Continuous ADMM for Transductive Inference in Higher-Order MRFs

CVPR 2018poster

This paper introduces a novel algorithm for transductive inference in higher-order MRFs, where the unary energies are parameterized by a variable classifier. The considered task is posed as a joint optimization problem in the continuous classifier parameters and the discrete label variables. In cont…

Cited by 10SourcePDFScholar
2018

Fight Ill-Posedness With Ill-Posedness: Single-Shot Variational Depth Super-Resolution From Shading

CVPR 2018poster

We put forward a principled variational approach for up-sampling a single depth map to the resolution of the companion color image provided by an RGB-D sensor. We combine heterogeneous depth and color data in order to jointly solve the ill-posed depth super-resolution and shape-from-shading problems…

2018

Incremental Semi-Supervised Learning from Streams for Object Classification

IROS 2018poster

The Label Propagation (LP) algorithm, first introduced by Zhu and Ghahramani [1], is a semi-supervised method used in transductive learning scenarios, where all data are available already in the beginning. In this work, we present a novel extension of the LP algorithm for applications where data sam…

Cited by 7SourceScholar
2018

MRF Optimization with Separable Convex Prior on Partially Ordered Labels

ECCV 2018poster

Solving a multi-labeling problem with a convex penalty can be achieved in polynomial time if the label set is totally ordered. In this paper we propose a generalization to partially ordered sets. To this end, we assume that the label set is the Cartesian product of totally ordered sets and the conve…

Cited by 3SourcePDFScholar
2018

Modular Vehicle Control for Transferring Semantic Information Between Weather Conditions Using GANs

CoRL 2018

Even though end-to-end supervised learning has shown promising results for sensorimotor control of self-driving cars, its performance is greatly affected by the weather conditions under which it was trained, showing poor generalization to unseen conditions. In this paper, we show how knowledge can b

2018

Omnidirectional DSO: Direct Sparse Odometry With Fisheye Cameras

RA-L 2018

We propose a novel real-time direct monocular visual odometry for omnidirectional cameras. Our method extends direct sparse odometry by using the unified omnidirectional model as a projection function, which can be applied to fisheye cameras with a field-of-view (FoV) well above 180°. This formulati

Cited by 99SourceScholar
2018

Online Photometric Calibration of Auto Exposure Video for Realtime Visual Odometry and SLAM

RA-L 2018

Recent direct visual odometry and SLAM algorithms have demonstrated impressive levels of precision. However, they require a photometric camera calibration in order to achieve competitive results. Hence, the respective algorithm cannot be directly applied to an off-the-shelf-camera or to a video sequ

Cited by 83SourceScholar
2018

StaticFusion: Background Reconstruction for Dense RGB-D SLAM in Dynamic Environments

ICRA 2018poster

Dynamic environments are challenging for visual SLAM as moving objects can impair camera pose tracking and cause corruptions to be integrated into the map. In this paper, we propose a method for robust dense RGB-D SLAM in dynamic environments which detects moving objects and simultaneously reconstru…

Cited by 248SourceScholar
2018

The TUM VI Benchmark for Evaluating Visual-Inertial Odometry

IROS 2018poster

Visual odometry and SLAM methods have a large variety of applications in domains such as augmented reality or robotics. Complementing vision sensors with inertial measurements tremendously improves tracking accuracy and robustness, and thus has spawned large interest in the development of visual-ine…

Cited by 520SourceScholar
2017

A Combinatorial Solution to Non-Rigid 3D Shape-To-Image Matching

CVPR 2017poster

We propose a combinatorial solution for the problem of non-rigidly matching a 3D shape to 3D image data. To this end, we model the shape as a triangular mesh and allow each triangle of this mesh to be rigidly transformed to achieve a suitable matching to the image. By penalising the distance and the…

Cited by 19PDFScholar
2017

A Non-Convex Variational Approach to Photometric Stereo Under Inaccurate Lighting

CVPR 2017poster

This paper tackles the photometric stereo problem in the presence of inaccurate lighting, obtained either by calibration or by an uncalibrated photometric stereo method. Based on a precise modeling of noise and outliers, a robust variational approach is introduced. It explicitly accounts for self-sh…

Cited by 73PDFScholar
2017

An Efficient Background Term for 3D Reconstruction and Tracking With Smooth Surface Models

CVPR 2017poster

We present a novel strategy to shrink and constrain a 3D model, represented as a smooth spline-like surface, within the visual hull of an object observed from one or multiple views. This new 'background' or 'silhouette' term combines the efficiency of previous approaches based on an image-plane dist…

Cited by 7PDFScholar
2017

De-noising, stabilizing and completing 3D reconstructions on-the-go using plane priors

ICRA 2017poster

Creating 3D maps on robots and other mobile devices has become a reality in recent years. Online 3D reconstruction enables many exciting applications in robotics and AR/VR gaming. However, the reconstructions are noisy and generally incomplete. Moreover, during online reconstruction, the surface cha…

Cited by 45SourceScholar
2017

Fast odometry and scene flow from RGB-D cameras based on geometric clustering

ICRA 2017poster

In this paper we propose an efficient solution to jointly estimate the camera motion and a piecewise-rigid scene flow from an RGB-D sequence. The key idea is to perform a two-fold segmentation of the scene, dividing it into geometric clusters that are, in turn, classified as static or moving element…

Cited by 142SourceScholar
2017

Image-Based Localization Using LSTMs for Structured Feature Correlation

ICCV 2017poster

In this work we propose a new CNN+LSTM architecture for camera pose regression for indoor and outdoor scenes. CNNs allow us to learn suitable feature representations for localization that are robust against motion blur and illumination changes. We make use of LSTM units on the CNN output, which play…

Cited by 658PDFScholar
2017

Intrinsic3D: High-Quality 3D Reconstruction by Joint Appearance and Geometry Optimization With Spatially-Varying Lighting

ICCV 2017poster

We introduce a novel method to obtain high-quality 3D reconstructions from consumer RGB-D sensors. Our core idea is to simultaneously optimize for geometry encoded in a signed distance field (SDF), textures from automatically-selected keyframes, and their camera poses along with material and scene l…

Cited by 136PDFScholar
2017

KillingFusion: Non-Rigid 3D Reconstruction Without Correspondences

CVPR 2017spotlight

We introduce a geometry-driven approach for real-time 3D reconstruction of deforming surfaces from a single RGB-D stream without any templates or shape priors. To this end, we tackle the problem of non-rigid registration by level set evolution without explicit correspondence search. Given a pair of…

Cited by 210PDFScholar
2017

Learning Proximal Operators: Using Denoising Networks for Regularizing Inverse Imaging Problems

ICCV 2017poster

While variational methods have been among the most powerful tools for solving linear inverse problems in imaging, deep (convolutional) neural networks have recently taken the lead in many challenging benchmarks. A remaining drawback of deep learning approaches is their requirement for an expensive r…

Cited by 443PDFcodeScholar
2017

Learning by Association -- A Versatile Semi-Supervised Training Method for Neural Networks

CVPR 2017poster

In many real-world scenarios, labeled data for a specific machine learning task is costly to obtain. Semi-supervised training methods make use of abundantly available unlabeled data and a smaller number of labeled examples. We propose a new framework for semi-supervised training of deep neural netwo…

Cited by 153PDFScholar
2017

Multi-view deep learning for consistent semantic mapping with RGB-D cameras

IROS 2017poster

Visual scene understanding is an important capability that enables robots to purposefully act in their environment. In this paper, we propose a novel deep neural network approach to predict semantic segmentation from RGB-D sequences. The key innovation is to train our network to predict multi-view c…

Cited by 175SourceScholar
2017

One-Shot Video Object Segmentation

CVPR 2017poster

This paper tackles the task of semi-supervised video object segmentation, i.e., the separation of an object from the background in a video, given the mask of the first frame. We present One-Shot Video Object Segmentation (OSVOS), based on a fully-convolutional neural network architecture that is abl…

Cited by 1167PDFScholar
2017

Product Manifold Filter: Non-Rigid Shape Correspondence via Kernel Density Estimation in the Product Space

CVPR 2017poster

Many algorithms for the computation of correspondences between deformable shapes rely on some variant of nearest neighbor matching in a descriptor space. Such are, for example, various point-wise correspondence recovery algorithms used as a post-processing stage in the functional correspondence fram…

Cited by 147PDFScholar
2017

Real-time trajectory replanning for MAVs using uniform B-splines and a 3D circular buffer

IROS 2017poster

In this paper, we present a real-time approach to local trajectory replanning for microaerial vehicles (MAVs). Current trajectory generation methods for multicopters achieve high success rates in cluttered environments, but assume that the environment is static and require prior knowledge of the map…

Cited by 249SourceScholar
2016

A Large Dataset to Train Convolutional Networks for Disparity, Optical Flow, and Scene Flow Estimation

CVPR 2016poster

Recent work has shown that optical flow estimation can be formulated as a supervised learning task and can be successfully solved with convolutional networks. Training of the so-called FlowNet was enabled by a large synthetically generated dataset. The present paper extends the concept of optical f…

Cited by 3436PDFScholar
2016

Efficient Globally Optimal 2D-To-3D Deformable Shape Matching

CVPR 2016poster

We propose the first algorithm for non-rigid 2D-to-3D shape matching, where the input is a 2D query shape as well as a 3D target shape and the output is a continuous matching curve represented as a closed contour on the 3D shape. We cast the problem as finding the shortest circular path on the produ…

Cited by 40PDFScholar
2016

Protein contact prediction from amino acid co-evolution using convolutional networks for graph-valued images

NeurIPS 2016oral

Proteins are the "building blocks of life", the most abundant organic molecules, and the central focus of most areas of biomedicine. Protein structure is strongly related to protein function, thus structure prediction is a crucial task on the way to solve many biological questions. A contact map is…

Cited by 53SourcePDFScholar
2016

Stream-based Active Learning for efficient and adaptive classification of 3D objects

ICRA 2016poster

We present a new Active Learning approach for classifying objects from streams of 3D point cloud data. The major problems here are the non-uniform occurrence of class instances and the unbalanced numbers of samples per class. We show that standard online learning methods based on decision trees perf…

Cited by 39SourceScholar
2016

Sublabel-Accurate Relaxation of Nonconvex Energies

CVPR 2016oral

We propose a novel spatially continuous framework for convex relaxations based on functional lifting. Our method can be interpreted as a sublabel-accurate solution to multilabel problems. We show that previously proposed functional lifting methods optimize an energy which is linear between two label…

Cited by 50PDFcodeScholar
2015

A primal-dual framework for real-time dense RGB-D scene flow

ICRA 2015poster

This paper presents the first method to compute dense scene flow in real-time for RGB-D cameras. It is based on a variational formulation where brightness constancy and geometric consistency are imposed. Accounting for the depth data provided by RGB-D cameras, regularization of the flow field is imp…

Cited by 151SourceScholar
2015

Adopting an Unconstrained Ray Model in Light-Field Cameras for 3D Shape Reconstruction

CVPR 2015poster

Due to their recent availability as off-the-shelf commercial devices, light-field cameras has gathered increasing attention from both scientific community and industrial operators. However, their composite imaging formation process hinders the ability to exploit the well consolidated stack of calibr…

Cited by 22SourcePDFScholar
2015

Dense Continuous-Time Tracking and Mapping With Rolling Shutter RGB-D Cameras

ICCV 2015poster

We propose a dense continuous-time tracking and mapping method for RGB-D cameras. We parametrize the camera trajectory using continuous B-splines and optimize the trajectory through dense, direct image alignment. Our method also directly models rolling shutter in both RGB and depth images within the…

Cited by 107PDFScholar
2015

Entropy Minimization for Convex Relaxation Approaches

ICCV 2015poster

Despite their enormous success in solving hard combinatorial problems, convex relaxation approaches often suffer from the fact that the computed solutions are far from binary and that subsequent heuristic binarization may substantially degrade the quality of computed solutions. In this paper, we pr…

Cited by 7PDFScholar
2015

FlowNet: Learning Optical Flow With Convolutional Networks

ICCV 2015poster

Convolutional neural networks (CNNs) have recently been very successful in a variety of computer vision tasks, especially on those linked to recognition. Optical flow estimation has not been among the tasks CNNs succeeded at. In this paper we construct CNNs which are capable of solving the optical f…

Cited by 4909PDFScholar
2015

Learning Nonlinear Spectral Filters for Color Image Reconstruction

ICCV 2015poster

This paper presents the idea of learning optimal filters for color image reconstruction based on a novel concept of nonlinear spectral image decompositions recently proposed by Guy Gilboa. We use a multiscale image decomposition approach based on total variation regularization and Bregman iterations…

Cited by 17PDFScholar
2015

Model-Based Tracking at 300Hz Using Raw Time-of-Flight Observations

ICCV 2015poster

Consumer depth cameras have dramatically improved our ability to track rigid, articulated, and deformable 3D objects in real-time. However, depth cameras have a limited temporal resolution (frame-rate) that restricts the accuracy and robustness of tracking, especially for fast or unpredictable motio…

Cited by 23PDFScholar
2015

Semi-supervised online learning for efficient classification of objects in 3D data streams

IROS 2015poster

We present a novel learning algorithm especially designed for challenging, large-scale classification problems in mobile robotics. Our method addresses two important aims: first it reduces the required amount of interaction with a human supervisor, which increases the level of autonomy of the learni…

Cited by 11SourceScholar