← Search

Federico Tombari

176 accepted papers

2026

3D-LATTE: Latent Space 3D Editing from Textual Instructions

CVPR 2026

Despite the recent success of multi-view diffusion models for text/image-based 3D asset generation, instruction-based editing of 3D assets lacks surprisingly far behind the quality of generation models. The main reason is that recent approaches using 2D priors suffer from view-inconsistent editing s

Cited by 0SourcecodeScholar
2026

AnyUp: Universal Feature Upsampling

ICLR 2026oral

We introduce AnyUp, a method for feature upsampling that can be applied to any vision feature at any resolution, without encoder-specific training. Existing learning-based upsamplers for features like DINO or CLIP need to be re-trained for every feature extractor and thus do not generalize to differ…

Cited by 0SourcecodeScholar
2026

DynaTok: Token-Based 4D Reconstruction from Partial Point Clouds

ICML 2026poster

We address the problem of 4D reconstruction from partial point cloud sequences, where observations from depth sensors are incomplete, unordered, and lack explicit point correspondence over time. Recovering coherent 4D geometry in this geometry-only setting is challenging due to missing observations …

Cited by 0SourceScholar
2026

Elastic3D: Controllable Stereo Video Conversion with Guided Latent Decoding

CVPR 2026

The growing demand for immersive 3D content calls for automated monocular-to-stereo video conversion. We present a controllable, direct end-to-end method for upgrading a conventional video to a binocular one. Our approach, based on (conditional) latent diffusion, avoids artifacts due to explicit dep

Cited by 0SourcecodeScholar
2026

Featurising Pixels from Dynamic 3D Scenes with Linear In-Context Learners

CVPR 2026

One of the most exciting applications of vision models involve pixel-level reasoning. Despite the abundance of vision foundation models, we still lack representations that effectively embed spatio-temporal properties of visual scenes at the pixel level. Existing frameworks either train on image-base

Cited by 0SourceScholar
2026

Future Dynamic 3D Reconstruction: A 3D World Model with Disentangled Ego-Motion

ICML 2026poster

Forecasting the evolution of dynamic environments is crucial for autonomous agents. While generative world models have recently achieved high photorealism in 2D video synthesis, by mixing within the image plane ego-motion and environmental dynamics, they exhibit physical inconsistencies, such as mor…

Cited by 0SourceScholar
2026

InstantPose: Zero-Shot Instance-Level 6D Pose Estimation from a Single View

ICRA 2026poster

Object pose estimation using visual data is crucial for robotic interaction with the environment. Many existing instance-level methods are restricted by their requirements for 3D CAD models or multiple object views, which limits their flexibility and generalizability. Overcoming this limitation is c…

Cited by 0SourceScholar
2026

LaRI: Layered Ray Intersections for Single-view 3D Geometric Reasoning

ICML 2026poster

We present Layered Ray Intersections (LaRI), a fully supervised method for occluded geometry reasoning from a single image. Unlike conventional depth estimation, which is limited to visible surfaces, LaRI predicts multiple surfaces intersected by the camera rays using layered point maps. Compared to…

Cited by 0SourceScholar
2026

OVI-MAP: Open-Vocabulary Instance-Semantic Mapping

CVPR 2026

Incremental open-vocabulary 3D instance-semantic mapping is essential for autonomous agents operating in complex everyday environments. However, it remains challenging due to the need for robust instance segmentation, real-time processing, and flexible open-set reasoning. Existing methods often rely

Cited by 0SourcecodeScholar
2026

PointDiT: Pixel-Space Diffusion for Monocular Geometry Estimation

ICML 2026poster

State-of-the-art single-image 3D reconstruction methods often rely on complex hybrid architectures or necessitate compressing geometry into latent spaces to leverage pre-trained latent diffusion models. In this work, we demonstrate that such architectural overhead is unnecessary. We introduce a mini…

Cited by 0SourceScholar
2026

RiemanLine: Riemannian Manifold Representation of 3D Lines for Factor Graph Optimization

AAAI 2026technical

Minimal parametrization of 3D lines plays a critical role in camera localization and structural mapping. Existing representations in robotics and computer vision predominantly handle independent lines, overlooking structural regularities such as sets of parallel lines that are pervasive in man-made

Cited by 0SourcePDFScholar
2026

SG2Loc: Sequential Visual Localization on 3D Scene Graphs

ICML 2026poster

Visual localization in complex environments remains a critical challenge for robotics and AR applications. Sequential localization, where pose estimates are refined over time, is important for autonomous agents. However, traditional methods often require storing extensive image databases or point cl…

Cited by 0SourceScholar
2026

Search3D: Hierarchical Open-Vocabulary 3D Segmentation

ICRA 2026poster

Open-vocabulary 3D segmentation enables the exploration of 3D spaces using free-form text descriptions. Existing methods for open-vocabulary 3D instance segmentation primarily focus on identifying object-level instances in a scene. However, they face challenges when it comes to understanding more fi…

2026

Text-to-3D by Stitching a Multi-view Reconstruction Network to a Video Generator

ICLR 2026oral

The rapid progress of large, pretrained models for both visual content generation and 3D reconstruction opens up new possibilities for text-to-3D generation. Intuitively, one could obtain a formidable 3D scene generator if one were able to combine the power of a modern latent text-to-video model as…

Cited by 0SourcecodeScholar
2026

The Art of Interrogation: Consistency Amplifies Factuality in Spatial Reasoning

ICML 2026poster

Current Large Reasoning Models (LRMs) exhibit remarkable general capabilities but significantly underperform in spatial reasoning tasks. Existing approaches treat this gap as a knowledge deficit, relying on supervised fine-tuning (SFT) to ingest labeled data from external vision sources or synthetic…

Cited by 0SourceScholar
2026

Understanding, Accelerating, and Improving MeanFlow Training

CVPR 2026

MeanFlow promises high-quality generative modeling in few steps, by jointly learning instantaneous and average velocity fields. Yet, the underlying training dynamics remain unclear. We analyze the interaction between the two velocities and find: (i) well-established instantaneous velocity is a prere

Cited by 0SourcecodeScholar
2025

Active Data Curation Effectively Distills Large-Scale Multimodal Models

CVPR 2025poster

Knowledge distillation (KD) is the de facto standard for compressing large-scale models into smaller ones. Prior works have explored ever more complex KD strategies involving different objective functions, teacher-ensembles, and weight inheritance. In this work we explore an alternative, yet simple…

Cited by 6SourcePDFScholar
2025

Contrastive Test-Time Composition of Multiple LoRA Models for Image Generation

ICCV 2025poster

Low-Rank Adaptation (LoRA) has emerged as a powerful and popular technique for personalization, enabling efficient adaptation of pre-trained image generation models for specific tasks without comprehensive retraining. While employing individual pre-trained LoRA models excels at representing single c…

Cited by 0SourcePDFScholar
2025

CubeDiff: Repurposing Diffusion-Based Image Models for Panorama Generation

ICLR 2025spotlight

We introduce a novel method for generating 360° panoramas from text prompts or images. Our approach leverages recent advances in 3D generation by employing multi-view diffusion models to jointly synthesize the six faces of a cubemap. Unlike previous methods that rely on processing equirectangular pr…

Cited by 3SourcePDFScholar
2025

ESCAPE: Equivariant Shape Completion via Anchor Point Encoding

CVPR 2025poster

Shape completion, a crucial task in 3D computer vision, involves predicting and filling the missing regions of scanned or partially observed objects. Current methods expect known pose or canonical coordinates and do not perform well under varying rotations, limiting their real-world applicability. W…

2025

Gatekeeper: Improving Model Cascades Through Confidence Tuning

NeurIPS 2025poster

Large-scale machine learning models deliver strong performance across a wide range of tasks but come with significant computational and resource constraints. To mitigate these challenges, local smaller models are often deployed alongside larger models, relying on routing and deferral mechanisms to o…

Cited by 0SourceScholar
2025

InstantPose: Zero-Shot Instance-Level 6D Pose Estimation From a Single View

RA-L 2025

Object pose estimation using visual data is crucial for robotic interaction with the environment. Many existing instance-level methods are restricted by their requirements for 3D CAD models or multiple object views, which limits their flexibility and generalizability. Overcoming this limitation is c

Cited by 4SourceScholar
2025

LODGE: Level-of-Detail Large-Scale Gaussian Splatting with Efficient Rendering

NeurIPS 2025spotlight

In this work, we present a novel level-of-detail (LOD) method for 3D Gaussian Splatting that enables real-time rendering of large-scale scenes on memory-constrained devices. Our approach introduces a hierarchical LOD representation that iteratively selects optimal subsets of Gaussians based on camer…

Cited by 0SourceScholar
2025

LayoutVLM: Differentiable Optimization of 3D Layout via Vision-Language Models

CVPR 2025poster

Spatial reasoning is a fundamental aspect of human cognition, enabling intuitive understanding and manipulation of objects in three-dimensional space. While foundation models demonstrate remarkable performance on some benchmarks, they still struggle with 3D reasoning tasks like arranging objects in…

Cited by 10SourcePDFScholar
2025

Learning Neural Exposure Fields for View Synthesis

NeurIPS 2025poster

Recent advances in neural scene representations have led to unprecedented quality in 3D reconstruction and view synthesis. Despite achieving high-quality results for common benchmarks with curated data, outputs often degrade for data that contain per image variations such as strong exposure changes,…

Cited by 0SourceScholar
2025

Learning to Prompt with Text Only Supervision for Vision-Language Models

AAAI 2025technical

Foundational vision-language models like CLIP are emerging as a promising paradigm in vision due to their excellent generalization. However, adapting these models for downstream tasks while maintaining their generalization remains challenging. In literature, one branch of methods adapts CLIP by lear…

2025

LiLoc: Lifelong Localization Using Adaptive Submap Joining and Egocentric Factor Graph

ICRA 2025

This paper proposes a versatile graph-based lifelong localization framework using LiDAR, LiLoc, which enhances its timeliness by maintaining a single central session while improves the accuracy through multi-modal factors between the central and subsidiary sessions. First, an adaptive submap joining

Cited by 3SourcecodeScholar
2025

LoRACLR: Contrastive Adaptation for Customization of Diffusion Models

CVPR 2025poster

Recent advances in text-to-image customization have enabled high-fidelity, context-rich generation of personalized images, allowing specific concepts to appear in a variety of scenarios. However, current methods struggle with combining multiple personalized models, often leading to attribute entangl…

Cited by 0SourcePDFScholar
2025

MOBIUS: Big-to-Mobile Universal Instance Segmentation via Multi-modal Bottleneck Fusion and Calibrated Decoder Pruning

ICCV 2025poster

Scaling up model size and training data has advanced foundation models for instance-level perception, achieving state-of-the-art in-domain and zero-shot performance across object detection and segmentation. However, their high computational cost limits adoption on resource-constrained platforms. We…

2025

Object-X: Learning to Reconstruct Multi-Modal 3D Object Representations

NeurIPS 2025poster

Learning effective multi-modal 3D representations of objects is essential for numerous applications, such as augmented reality and robotics. Existing methods often rely on task-specific embeddings that are tailored either for semantic understanding or geometric reconstruction. As a result, these e…

Cited by 0SourceScholar
2025

Omnia de EgoTempo: Benchmarking Temporal Understanding of Multi-Modal LLMs in Egocentric Videos

CVPR 2025poster

Understanding fine-grained temporal dynamics is crucial in egocentric videos, where continuous streams capture frequent, close-up interactions with objects. In this work, we bring to light that current egocentric video question-answering datasets often include questions that can be answered using on…

2025

One2Any: One-Reference 6D Pose Estimation for Any Object

CVPR 2025poster

6D object pose estimation remains challenging for many applications due to dependencies on complete 3D models, multi-view images, or training limited to specific object categories. These requirements make generalization to novel objects difficult for which neither 3D models nor multi-view images may…

2025

Prior2Former - Evidential Modeling of Mask Transformers for Assumption-Free Open-World Panoptic Segmentation

ICCV 2025poster

In panoptic segmentation, individual instances must be separated within semantic classes. As state-of-the-art methods rely on a pre-defined set of classes, they struggle with novel categories and out-of-distribution (OOD) data. This is particularly problematic in safety-critical applications, such a…

Cited by 0SourcePDFScholar
2025

RelationField: Relate Anything in Radiance Fields

CVPR 2025poster

Neural radiance fields are an emerging 3D scene representation and recently even been extended to learn features for scene understanding by distilling open-vocabulary features from vision-language models. However, current method primarily focus on object-centric representations, supporting object se…

2025

Search3D: Hierarchical Open-Vocabulary 3D Segmentation

RA-L 2025

Open-vocabulary 3D segmentation enables exploration of 3D spaces using free-form text descriptions. Existing methods for open-vocabulary 3D instance segmentation primarily focus on identifying <italic xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">object</i

Cited by 31SourceScholar
2025

Semantic Library Adaptation: LoRA Retrieval and Fusion for Open-Vocabulary Semantic Segmentation

CVPR 2025poster

Open-vocabulary semantic segmentation models associate vision and text to label pixels from an undefined set of classes using textual queries, providing versatile performance on novel datasets. However, large shifts between training and test domains degrade their performance, requiring fine-tuning f…

2025

Test-Time Visual In-Context Tuning

CVPR 2025poster

Visual in-context learning (VICL), as a new paradigm in computer vision, allows the model to rapidly adapt to various tasks with only a handful of prompts and examples. While effective, the existing VICL paradigm exhibits poor generalizability under distribution shifts. In this work, we propose test…

2025

TokenFormer: Rethinking Transformer Scaling with Tokenized Model Parameters

ICLR 2025spotlight

Transformers have become the predominant architecture in foundation models due to their excellent performance across various domains. However, the substantial cost of scaling these models remains a significant concern. This problem arises primarily from their dependence on a fixed number of paramete…

2025

UIP2P: Unsupervised Instruction-based Image Editing via Edit Reversibility Constraint

ICCV 2025poster

We propose an unsupervised instruction-based image editing approach that removes the need for ground-truth edited images during training. Existing methods rely on supervised learning with triplets of input images, ground-truth edited images, and edit instructions. These triplets are typically genera…

Cited by 0SourcePDFScholar
2025

UNOPose: Unseen Object Pose Estimation with an Unposed RGB-D Reference Image

CVPR 2025poster

Unseen object pose estimation methods often rely on CAD models or multiple reference views, making the onboarding stage costly. To simplify reference acquisition, we aim to estimate the unseen object's pose through a single unposed RGB-D reference image. While previous works leverage reference image…

2025

Video Perception Models for 3D Scene Synthesis

NeurIPS 2025poster

Automating the expert-dependent and labor-intensive task of 3D scene synthesis would significantly benefit fields such as architectural design, robotics simulation, and virtual reality. Recent approaches to 3D scene synthesis often rely on the commonsense reasoning of large language models (LLMs) or…

Cited by 0SourceScholar
2024

BRAVE: Broadening the visual encoding of vision-language models

ECCV 2024oral

"Vision-language models (VLMs) are typically composed of a vision encoder, e.g. CLIP, and a language model (LM) that interprets the encoded features to solve downstream tasks. Despite remarkable progress, VLMs are subject to several shortcomings due to the limited capabilities of vision encoders, e.…

2024

CONFORM: Contrast is All You Need for High-Fidelity Text-to-Image Diffusion Models

CVPR 2024poster

Images produced by text-to-image diffusion models might not always faithfully represent the semantic intent of the provided text prompt where the model might overlook or entirely fail to produce certain objects. While recent studies propose various solutions they often require customly tailored func…

Cited by 21SourcePDFScholar
2024

Denoising Diffusion via Image-Based Rendering

ICLR 2024poster

Generating 3D scenes is a challenging open problem, which requires synthesizing plausible content that is fully consistent in 3D space. While recent methods such as neural radiance fields excel at view synthesis and 3D reconstruction, they cannot synthesize plausible details in unobserved regions si…

Cited by 11SourcePDFScholar
2024

Diffusion Bridges for 3D Point Cloud Denoising

ECCV 2024poster

"In this work, we address the task of point cloud denoising using a novel framework adapting Diffusion Schrödinger bridges to unstructured data like point sets. Unlike previous works that predict point-wise displacements from point features or learned noise distributions, our method learns an optim…

2024

EchoScene: Indoor Scene Generation via Information Echo over Scene Graph Diffusion

ECCV 2024poster

"We present EchoScene, an interactive and controllable generative model that generates 3D indoor scenes on scene graphs. EchoScene leverages a dual-branch diffusion model that dynamically adapts to scene graphs. Existing methods struggle to handle scene graphs due to varying numbers of nodes, multip…

Cited by 23SourcePDFScholar
2024

Extracting Training Data From Document-Based VQA Models

ICML 2024poster

Vision-Language Models (VLMs) have made remarkable progress in document-based Visual Question Answering (i.e., responding to queries about the contents of an input document provided as an image). In this work, we show these models can memorize responses for training samples and regurgitate them even…

Cited by 7SourcePDFScholar
2024

GeoGaussian: Geometry-aware Gaussian Splatting for Scene Rendering

ECCV 2024poster

"During the Gaussian Splatting optimization process, the scene geometry can gradually deteriorate if its structure is not deliberately preserved, especially in non-textured regions such as walls, ceilings, and furniture surfaces. This degradation significantly affects the rendering quality of novel…

Cited by 25SourcePDFScholar
2024

HyperSDFusion: Bridging Hierarchical Structures in Language and Geometry for Enhanced 3D Text2Shape Generation

CVPR 2024poster

3D shape generation from text is a fundamental task in 3D representation learning. The text-shape pairs exhibit a hierarchical structure where a general text like "chair" covers all 3D shapes of the chair while more detailed prompts refer to more specific shapes. Furthermore both text and 3D shapes…

2024

KP-RED: Exploiting Semantic Keypoints for Joint 3D Shape Retrieval and Deformation

CVPR 2024poster

In this paper we present KP-RED a unified KeyPoint-driven REtrieval and Deformation framework that takes object scans as input and jointly retrieves and deforms the most geometrically similar CAD models from a pre-processed database to tightly match the target. Unlike existing dense matching based m…

2024

Know Your Neighbors: Improving Single-View Reconstruction via Spatial Vision-Language Reasoning

CVPR 2024poster

Recovering the 3D scene geometry from a single view is a fundamental yet ill-posed problem in computer vision. While classical depth estimation methods infer only a 2.5D scene representation limited to the image plane recent approaches based on radiance fields reconstruct a full 3D representation. H…

2024

MOHO: Learning Single-view Hand-held Object Reconstruction with Multi-view Occlusion-Aware Supervision

CVPR 2024poster

Previous works concerning single-view hand-held object reconstruction typically rely on supervision from 3D ground-truth models which are hard to collect in real world. In contrast readily accessible hand-object videos offer a promising training data source but they only give heavily occluded object…

Cited by 11SourcePDFScholar
2024

NEWTON: Neural View-Centric Mapping for On-the-Fly Large-Scale SLAM

RA-L 2024

Neural field-based 3D representations have recently been adopted in many areas including SLAM systems. Current neural SLAM or online mapping systems lead to impressive results in the presence of simple captures, but they rely on a world-centric map representation as only a single neural field model

Cited by 25SourceScholar
2024

Open-Structure: Structural Benchmark Dataset for SLAM Algorithms

RA-L 2024

This letter presents Open-Structure, a novel benchmark dataset for evaluating visual odometry and SLAM methods. Compared to existing public datasets that primarily offer raw images, Open-Structure provides direct access to point and line measurements, correspondences, structural associations, and co

Cited by 5SourcecodeScholar
2024

OpenNeRF: Open Set 3D Neural Scene Segmentation with Pixel-Wise Features and Rendered Novel Views

ICLR 2024poster

Large visual-language models (VLMs), like CLIP, enable open-set image segmentation to segment arbitrary concepts from an image in a zero-shot manner. This goes beyond the traditional closed-set assumption, i.e., where models can only segment classes from a pre-defined training set. More recently, fi…

Cited by 33SourcePDFScholar
2024

PhysAvatar: Learning the Physics of Dressed 3D Avatars from Visual Observations

ECCV 2024poster

"[width=0.9]figure/teaserv 4.pdf Figure 1: PhysAvatar is a novel framework that captures the physics of dressed 3D avatars from visual observations, enabling a wide spectrum of applications, such as (a) animation, (b) relighting, and (c) redressing, with high-fidelity rendering results."

2024

Physics-Encoded Graph Neural Networks for Deformation Prediction under Contact

ICRA 2024poster

In robotics, it’s crucial to understand object deformation during tactile interactions. A precise understanding of deformation can elevate robotic simulations and have broad implications across different industries. We introduce a method using Physics-Encoded Graph Neural Networks (GNNs) for such pr…

Cited by 4SourceScholar
2024

SG-Bot: Object Rearrangement via Coarse-to-Fine Robotic Imagination on Scene Graphs

ICRA 2024poster

Object rearrangement is pivotal in robotic-environment interactions, representing a significant capability in embodied AI. In this paper, we present SG-Bot, a novel rearrangement framework that utilizes a coarse-to-fine scheme with a scene graph as the scene representation. Unlike previous methods t…

Cited by 25SourceScholar
2024

SILC: Improving Vision Language Pretraining with Self-Distillation

ECCV 2024poster

"Image-Text pretraining on web-scale image caption datasets has become the default recipe for open vocabulary classification and retrieval models thanks to the success of CLIP and its variants. Several works have also used CLIP features for dense prediction tasks and have shown the emergence of open…

Cited by 27SourcePDFScholar
2024

SceneFun3D: Fine-Grained Functionality and Affordance Understanding in 3D Scenes

CVPR 2024poster

Existing 3D scene understanding methods are heavily focused on 3D semantic and instance segmentation. However identifying objects and their parts only constitutes an intermediate step towards a more fine-grained goal which is effectively interacting with the functional interactive elements (e.g. han…

Cited by 35SourcePDFScholar
2024

SceneGraphLoc: Cross-Modal Coarse Visual Localization on 3D Scene Graphs

ECCV 2024poster

"We introduce the task of localizing an input image within a multi-modal reference map represented by a collection of 3D scene graphs. These scene graphs comprise multiple modalities, including object-level point clouds, images, attributes, and relationships between objects, offering a lightweight a…

2024

SecondPose: SE(3)-Consistent Dual-Stream Feature Fusion for Category-Level Pose Estimation

CVPR 2024poster

Category-level object pose estimation aiming to predict the 6D pose and 3D size of objects from known categories typically struggles with large intra-class shape variation. Existing works utilizing mean shapes often fall short of capturing this variation. To address this issue we present SecondPose…

2024

Segment3D: Learning Fine-Grained Class-Agnostic 3D Segmentation without Manual Labels

ECCV 2024poster

"Current 3D scene segmentation methods are heavily dependent on manually annotated 3D training datasets. Such manual annotations are labor-intensive, and often lack fine-grained details. Furthermore, models trained on this data typically struggle to recognize object classes beyond the annotated trai…

Cited by 34SourcePDFScholar
2024

Self-supervised Shape Completion via Involution and Implicit Correspondences

ECCV 2024poster

"3D shape completion is traditionally solved using supervised training or by distribution learning on complete shape examples. Recently self-supervised learning approaches that do not require any complete 3D shape examples have gained more interests. In this paper, we propose a non-adversarial self-…

2024

SemiVL: Semi-Supervised Semantic Segmentation with Vision-Language Guidance

ECCV 2024poster

"In semi-supervised semantic segmentation, a model is trained with a limited number of labeled images along with a large corpus of unlabeled images to reduce the high annotation effort. While previous methods are able to learn good segmentation boundaries, they are prone to confuse classes with simi…

2024

Stylebreeder: Exploring and Democratizing Artistic Styles through Text-to-Image Models

NeurIPS 2024poster

Text-to-image models are becoming increasingly popular, revolutionizing the landscape of digital art creation by enabling highly detailed and creative visual content generation. These models have been widely employed across various domains, particularly in art generation, where they facilitate a bro…

2024

Text-Conditioned Resampler For Long Form Video Understanding

ECCV 2024poster

"In this paper we present a text-conditioned video resampler (TCR) module that uses a pre-trained and frozen visual encoder and large language model (LLM) to process long video sequences for a task. TCR localises relevant visual features from the video given a text condition and provides them to a L…

Cited by 12SourcePDFScholar
2024

UniSDF: Unifying Neural Representations for High-Fidelity 3D Reconstruction of Complex Scenes with Reflections

NeurIPS 2024poster

Neural 3D scene representations have shown great potential for 3D reconstruction from 2D images. However, reconstructing real-world captures of complex scenes still remains a challenge. Existing generic 3D reconstruction methods often struggle to represent fine geometric details and do not adequatel…

2024

Zero123-6D: Zero-shot Novel View Synthesis for RGB Category-level 6D Pose Estimation

IROS 2024

Estimating the pose of objects through vision is essential to make robotic platforms interact with the environment. Yet, it presents many challenges, often related to the lack of flexibility and generalizability of state-of-the-art solutions. Diffusion models are a cutting-edge neural architecture t

Cited by 13SourceScholar
2023

CommonScenes: Generating Commonsense 3D Indoor Scenes with Scene Graph Diffusion

NeurIPS 2023poster

Controllable scene synthesis aims to create interactive environments for numerous industrial use cases. Scene graphs provide a highly suitable interface to facilitate these applications by abstracting the scene context in a compact manner. Existing methods, reliant on retrieval from extensive databa…

2023

DDF-HO: Hand-Held Object Reconstruction via Conditional Directed Distance Field

NeurIPS 2023poster

Reconstructing hand-held objects from a single RGB image is an important and challenging problem. Existing works utilizing Signed Distance Fields (SDF) reveal limitations in comprehensively capturing the complex hand-object interactions, since SDF is only reliable within the proximity of the target…

2023

Dynamic Hyperbolic Attention Network for Fine Hand-object Reconstruction

ICCV 2023poster

Reconstructing both objects and hands in 3D from a single RGB image is complex. Existing methods rely on manually defined hand-object constraints in Euclidean space, leading to suboptimal feature learning. Compared with Euclidean space, hyperbolic space better preserves the geometric properties of m…

Cited by 14PDFScholar
2023

I2MVFormer: Large Language Model Generated Multi-View Document Supervision for Zero-Shot Image Classification

CVPR 2023highlight

Recent works have shown that unstructured text (documents) from online sources can serve as useful auxiliary information for zero-shot image classification. However, these methods require access to a high-quality source like Wikipedia and are limited to a single source of information. Large Language…

2023

IPCC-TP: Utilizing Incremental Pearson Correlation Coefficient for Joint Multi-Agent Trajectory Prediction

CVPR 2023poster

Reliable multi-agent trajectory prediction is crucial for the safe planning and control of autonomous systems. Compared with single-agent cases, the major challenge in simultaneously processing multiple agents lies in modeling complex social interactions caused by various driving intentions and road…

Cited by 19SourcePDFScholar
2023

Incremental 3D Semantic Scene Graph Prediction From RGB Sequences

CVPR 2023poster

3D semantic scene graphs are a powerful holistic representation as they describe the individual objects and depict the relation between them. They are compact high-level graphs that enable many tasks requiring scene reasoning. In real-world settings, existing 3D estimation methods produce robust pre…

2023

Introducing Language Guidance in Prompt-based Continual Learning

ICCV 2023poster

Continual Learning aims to learn a single model on a sequence of tasks without having access to data from previous tasks. The biggest challenge in the domain still remains catastrophic forgetting: a loss in performance on seen classes of earlier tasks. Some existing methods rely on an expensive repl…

Cited by 56PDFcodeScholar
2023

MonoGraspNet: 6-DoF Grasping with a Single RGB Image

ICRA 2023poster

6-DoF robotic grasping is a long-lasting but un-solved problem. Recent methods utilize strong 3D networks to extract geometric grasping representations from depth sensors, demonstrating superior accuracy on common objects but performing unsatisfactorily on photometrically challenging objects, e.g.,…

Cited by 39SourceScholar
2023

OPA-3D: Occlusion-Aware Pixel-Wise Aggregation for Monocular 3D Object Detection

RA-L 2023

Monocular 3D object detection has recently made a significant leap forward thanks to the use of pre-trained depth estimators for pseudo-LiDAR recovery. Yet, such two-stage methods typically suffer from overfitting and are incapable of explicitly encapsulating the geometric relation between depth and

Cited by 39SourceScholar
2023

OpenMask3D: Open-Vocabulary 3D Instance Segmentation

NeurIPS 2023poster

We introduce the task of open-vocabulary 3D instance segmentation. Current approaches for 3D instance segmentation can typically only recognize object categories from a pre-defined closed set of classes that are annotated in the training datasets. This results in important limitations for real-world…

2023

Robust Monocular Depth Estimation under Challenging Conditions

ICCV 2023accepted

While state-of-the-art monocular depth estimation approaches achieve impressive results in ideal settings, they are highly unreliable under challenging illumination and weather conditions, such as at nighttime or in the presence of rain. In this paper, we uncover these safety-critical issues and tac…

2023

SPARF: Neural Radiance Fields From Sparse and Noisy Poses

CVPR 2023highlight

Neural Radiance Field (NeRF) has recently emerged as a powerful representation to synthesize photorealistic novel views. While showing impressive performance, it relies on the availability of dense input views with highly accurate camera poses, thus limiting its application in real-world scenarios.…

2023

Segmenting Known Objects and Unseen Unknowns without Prior Knowledge

ICCV 2023poster

Panoptic segmentation methods assign a known class to each pixel given in input. Even for state-of-the-art approaches, this inevitably enforces decisions that systematically lead to wrong predictions for objects outside the training categories. However, robustness against out-of-distribution samples…

Cited by 10PDFcodeScholar
2023

Self-Supervised Category-Level 6D Object Pose Estimation With Optical Flow Consistency

RA-L 2023

Category-level 6D object pose estimation aims at determining the pose of an object of a given category. Most current state-of-the-art methods require a significant amount of real training data to supervise their models. Moreover, annotating the 6D pose is very time consuming, error-prone, and it doe

Cited by 19SourceScholar
2023

Shape, Pose, and Appearance From a Single Image via Bootstrapped Radiance Field Inversion

CVPR 2023poster

Neural Radiance Fields (NeRF) coupled with GANs represent a promising direction in the area of 3D reconstruction from a single view, owing to their ability to efficiently model arbitrary topologies. Recent work in this area, however, has mostly focused on synthetic datasets where exact ground-truth…

2023

Socratic Models: Composing Zero-Shot Multimodal Reasoning with Language

ICLR 2023top-25%

We investigate how multimodal prompt engineering can use language as the intermediate representation to combine complementary knowledge from different pretrained (potentially multimodal) language models for a variety of tasks. This approach is both distinct from and complementary to the dominant par…

2023

SparseFusion: Fusing Multi-Modal Sparse Representations for Multi-Sensor 3D Object Detection

ICCV 2023poster

By identifying four important components of existing LiDAR-camera 3D object detection methods (LiDAR and camera candidates, transformation, and fusion outputs), we observe that all existing methods either find dense candidates or yield dense representations of scenes. However, given that objects occ…

Cited by 77PDFcodeScholar
2023

SupeRGB-D: Zero-Shot Instance Segmentation in Cluttered Indoor Environments

RA-L 2023

Object instance segmentation is a key challenge for indoor robots navigating cluttered environments with many small objects. Limitations in 3D sensing capabilities often make it difficult to detect every possible object. While deep learning approaches may be effective for this problem, manually anno

Cited by 14SourcecodeScholar
2023

Towards Long-Term Retrieval-Based Visual Localization in Indoor Environments With Changes

RA-L 2023

Visual localization is a challenging task due to the presence of illumination changes, occlusion, and perception from novel viewpoints. Re-localizing the camera pose in long-term setups raises difficulties caused by changes in scene appearance and geometry introduced by human or natural deterioratio

Cited by 12SourceScholar
2023

U-RED: Unsupervised 3D Shape Retrieval and Deformation for Partial Point Clouds

ICCV 2023poster

In this paper, we propose U-RED, an Unsupervised shape REtrieval and Deformation pipeline that takes an arbitrary object observation as input, typically captured by RGB images or scans, and jointly retrieves and deforms the geometrically similar CAD models from a pre-established database to tightly…

Cited by 3PDFcodeScholar
2022

3D Compositional Zero-Shot Learning with DeCompositional Consensus

ECCV 2022poster

"Parts represent a basic unit of geometric and semantic similarity across different objects. We argue that part knowledge should be composable beyond the observed object classes. Towards this, we present 3D Compositional Zero-shot Learning as a problem of part generalization from seen to unseen obje…

Cited by 15SourcePDFScholar
2022

3D-VField: Adversarial Augmentation of Point Clouds for Domain Generalization in 3D Object Detection

CVPR 2022poster

As 3D object detection on point clouds relies on the geometrical relationships between the points, non-standard object shapes can hinder a method's detection capability. However, in safety-critical settings, robustness to out-of-domain and long-tail samples is fundamental to circumvent dangerous iss…

Cited by 69PDFcodeScholar
2022

Bending Graphs: Hierarchical Shape Matching Using Gated Optimal Transport

CVPR 2022poster

Shape matching has been a long-studied problem for the computer graphics and vision community. The objective is to predict a dense correspondence between meshes that have a certain degree of deformation. Existing methods either consider the local description of sampled points or discover corresponde…

Cited by 24PDFcodeScholar
2022

CertainNet: Sampling-Free Uncertainty Estimation for Object Detection

RA-L 2022

Estimating the uncertainty of a neural network plays a fundamental role in safety-critical settings. In perception for autonomous driving, measuring the uncertainty means providing additional calibrated information to downstream tasks, such as path planning, that can use it towards safe navigation.

Cited by 29SourceScholar
2022

CloudAttention: Efficient Multi-Scale Attention Scheme For 3D Point Cloud Learning

IROS 2022poster

Processing 3D data efficiently has always been a challenge. Spatial operations on large-scale point clouds, stored as sparse data, require extra cost. Attracted by the success of transformers, researchers are using multi-head attention for vision tasks. However, attention calculations in transformer…

Cited by 5SourcecodeScholar
2022

GOCA: Guided Online Cluster Assignment for Self-Supervised Video Representation Learning

ECCV 2022poster

"Clustering is a ubiquitous tool in unsupervised learning. Most of the existing self-supervised representation learning methods typically cluster samples based on visually dominant features. While this works well for image-based selfsupervision, it often fails for videos, which require understanding…

2022

GPV-Pose: Category-Level Object Pose Estimation via Geometry-Guided Point-Wise Voting

CVPR 2022poster

While 6D object pose estimation has recently made a huge leap forward, most methods can still only handle a single or a handful of different objects, which limits their applications. To circumvent this problem, category-level object pose estimation has recently been revamped, which aims at predictin…

Cited by 147PDFcodeScholar
2022

I2DFormer: Learning Image to Document Attention for Zero-Shot Image Classification

NeurIPS 2022accept

Despite the tremendous progress in zero-shot learning (ZSL), the majority of existing methods still rely on human-annotated attributes, which are difficult to annotate and scale. An unsupervised alternative is to represent each class using the word embedding associated with its semantic class name.…

Cited by 55SourcePDFScholar
2022

Implicit Neural Representations for Image Compression

ECCV 2022poster

"Implicit Neural Representations (INRs) gained attention as a novel and effective representation for various data types. Recently, prior work applied INRs to image compressing. Such compression algorithms are promising candidates as a general purpose approach for any coordinate-based data modality.…

2022

Object-Aware Monocular Depth Prediction With Instance Convolutions

RA-L 2022

With the advent of deep learning, estimating depth from a single RGB image has recently received a lot of attention, being capable of empowering many different applications ranging from path planning for robotics to computational cinematography. Nevertheless,while the depth maps are in their entiret

Cited by 3SourcecodeScholar
2022

On the Practicality of Deterministic Epistemic Uncertainty

ICML 2022spotlight

A set of novel approaches for estimating epistemic uncertainty in deep neural networks with a single forward pass has recently emerged as a valid alternative to Bayesian Neural Networks. On the premise of informative representations, these deterministic uncertainty methods (DUMs) achieve strong perf…

2022

RBP-Pose: Residual Bounding Box Projection for Category-Level Pose Estimation

ECCV 2022poster

"Category-level object pose estimation aims to predict the 6D pose as well as the 3D metric size of previously unseen objects from a known set of categories. Recent methods harness shape prior adaptation to map the observed point cloud into the canonical space and apply Umeyama’s algorithm to recove…

2022

SHIFT: A Synthetic Driving Dataset for Continuous Multi-Task Domain Adaptation

CVPR 2022poster

Adapting to a continuously evolving environment is a safety-critical challenge inevitably faced by all autonomous-driving systems. Existing image- and video-based driving datasets, however, fall short of capturing the mutable nature of the real world. In this paper, we introduce the largest syntheti…

Cited by 166PDFScholar
2022

SSP-Pose: Symmetry-Aware Shape Prior Deformation for Direct Category-Level Object Pose Estimation

IROS 2022poster

Category-level pose estimation is a challenging problem due to intra-class shape variations. Recent methods deform pre-computed shape priors to map the observed point cloud into the normalized object coordinate space and then retrieve the pose via post-processing, i.e., Umeyama's Algorithm. The shor…

Cited by 41SourceScholar
2022

Time-to-Label: Temporal Consistency for Self-Supervised Monocular 3D Object Detection

RA-L 2022

Monocular 3D object detection continues to attract attention due to the cost benefits and wider availability of RGB cameras. Despite the recent advances and the ability to acquire data at scale, annotation cost and complexity still limit the size of 3D object detection datasets in the supervised set

Cited by 9SourceScholar
2022

ZebraPose: Coarse To Fine Surface Encoding for 6DoF Object Pose Estimation

CVPR 2022poster

Establishing correspondences from image to 3D has been a key task of 6DoF object pose estimation for a long time. To predict pose more accurately, deeply learned dense maps replaced sparse templates. Dense methods also improved pose estimation in the presence of occlusion. More recently researchers…

Cited by 176PDFcodeScholar
2021

Content Disentanglement for Semantically Consistent Synthetic-to-Real Domain Adaptation

IROS 2021poster

Synthetic data generation is an appealing approach to generate novel traffic scenarios in autonomous driving. However, deep learning perception algorithms trained solely on synthetic data encounter serious performance drops when they are tested on real data. Such performance drops are commonly attri…

Cited by 9SourcecodeScholar
2021

GDR-Net: Geometry-Guided Direct Regression Network for Monocular 6D Object Pose Estimation

CVPR 2021poster

6D pose estimation from a single RGB image is a fundamental task in computer vision. The current top-performing deep learning-based methods rely on an indirect strategy, i.e., first establishing 2D-3D correspondences between the coordinates in the image plane and object coordinate system, and then a…

Cited by 462PDFcodeScholar
2021

Graph-to-3D: End-to-End Generation and Manipulation of 3D Scenes Using Scene Graphs

ICCV 2021poster

Controllable scene synthesis consists of generating 3D information that satisfy underlying specifications. Thereby, these specifications should be abstract, i.e. allowing easy user interaction, whilst providing enough interface for detailed control. Scene graphs are representations of a scene, compo…

Cited by 85PDFcodeScholar
2021

Learning Graph Embeddings for Compositional Zero-Shot Learning

CVPR 2021poster

In compositional zero-shot learning, the goal is to recognize unseen compositions (e.g. old dog) of observed visual primitives states (e.g. old, cute) and objects (e.g. car, dog)in the training set. This is challenging because the same state can for example alter the visual appearance of a dog drast…

Cited by 199PDFcodeScholar
2021

ManhattanSLAM: Robust Planar Tracking and Mapping Leveraging Mixture of Manhattan Frames

ICRA 2021poster

In this paper, a robust RGB-D SLAM system is proposed to utilize the structural information in indoor scenes, allowing for accurate tracking and efficient dense mapping on a CPU. Prior works have used the Manhattan World (MW) assumption to estimate low-drift camera pose, in turn limiting the applica…

Cited by 85SourcecodeScholar
2021

Panoster: End-to-End Panoptic Segmentation of LiDAR Point Clouds

RA-L 2021

Panoptic segmentation has recently unified semantic and instance segmentation, previously addressed separately, thus taking a step further towards creating more comprehensive and efficient perception systems. In this letter, we present Panoster, a novel proposal-free panoptic segmentation method for

Cited by 76SourceScholar
2021

SO-Pose: Exploiting Self-Occlusion for Direct 6D Pose Estimation

ICCV 2021poster

Directly regressing all 6 degrees-of-freedom (6DoF) for the object pose (i.e. the 3D rotation and translation) in a cluttered environment from a single RGB image is a challenging problem. While end-to-end methods have recently demonstrated promising results at high efficiency, they are still inferio…

Cited by 160PDFcodeScholar
2021

SRH-Net: Stacked Recurrent Hourglass Network for Stereo Matching

RA-L 2021

The cost aggregation strategy shows a crucialrole in learning-based stereo matching tasks, where 3D convolutional filters obtain state of the art but require intensive computation resources, while 2D operations need less GPU memory but are sensitive to domain shift. In this letter, we decouple the 4

Cited by 13SourcecodeScholar
2021

SceneGraphFusion: Incremental 3D Scene Graph Prediction From RGB-D Sequences

CVPR 2021poster

Scene graphs are a compact and explicit representation successfully used in a variety of 2D scene understanding tasks. This work proposes a method to build up semantic scene graphs from a 3D environment incrementally given a sequence of RGB-D frames. To this end, we aggregate PointNet features from…

Cited by 186PDFScholar
2021

Semantic Image Alignment for Vehicle Localization

IROS 2021poster

Accurate and reliable localization is a fundamental requirement for autonomous vehicles to use map information in higher-level tasks such as navigation or planning. In this paper, we present a novel approach to vehicle localization in dense semantic maps, including vectorized high-definition maps or…

Cited by 6SourceScholar
2021

TSDF++: A Multi-Object Formulation for Dynamic Object Tracking and Reconstruction

ICRA 2021poster

The ability to simultaneously track and reconstruct multiple objects moving in the scene is of the utmost importance for robotic tasks such as autonomous navigation and interaction. Virtually all of the previous attempts to map multiple dynamic objects have evolved to store individual objects in sep…

Cited by 35SourcecodeScholar
2021

Unconditional Scene Graph Generation

ICCV 2021poster

Despite recent advancements in single-domain or single-object image generation, it is still challenging to generate complex scenes containing diverse, multiple objects and their interactions. Scene graphs, composed of nodes as objects and directed-edges as relationships among objects, offer an alter…

Cited by 35PDFScholar
2021

Unsupervised Traffic Scene Generation with Synthetic 3D Scene Graphs

IROS 2021poster

Image synthesis driven by computer graphics achieved recently a remarkable realism, yet synthetic image data generated this way reveals a significant domain gap with respect to real-world data. This is especially true in autonomous driving scenarios, which represent a critical aspect for over-coming…

Cited by 12SourceScholar
2020

Adversarial Appearance Learning in Augmented Cityscapes for Pedestrian Recognition in Autonomous Driving

ICRA 2020poster

In the autonomous driving area synthetic data is crucial for cover specific traffic scenarios which autonomous vehicle must handle. This data commonly introduces domain gap between synthetic and real domains. In this paper we deploy data augmentation to generate custom traffic scenarios with VRUs in…

Cited by 10SourceScholar
2020

Ambiguity in Sequential Data: Predicting Uncertain Futures With Recurrent Models

RA-L 2020

Ambiguity is inherently present in many machine learning tasks, but especially for sequential models seldom accounted for, as most only output a single prediction. In this work we propose an extension of the Multiple Hypothesis Prediction (MHP) model to handle ambiguous predictions with sequential d

Cited by 7SourceScholar
2020

Beyond Controlled Environments: 3D Camera Re-Localization in Changing Indoor Scenes

ECCV 2020poster

Long-term camera re-localization is an important task with numerous computer vision and robotics applications. Whilst various outdoor benchmarks exist that target lighting, weather and seasonal changes, far less attention has been paid to appearance changes that occur indoors. This has led to a mism…

2020

Binary DAD-Net: Binarized Driveable Area Detection Network for Autonomous Driving

ICRA 2020poster

Driveable area detection is a key component for various applications in the field of autonomous driving (AD), such as ground-plane detection, obstacle detection and maneuver planning. Additionally, bulky and over-parameterized networks can be easily forgone and replaced with smaller networks for fas…

Cited by 28SourceScholar
2020

Co-Planar Parametrization for Stereo-SLAM and Visual-Inertial Odometry

RA-L 2020

This letter proposes a novel SLAM framework for stereo and visual inertial odometry estimation. It builds an efficient and robust parametrization of co-planar points and lines which leverages specific geometric constraints to improve camera pose optimization in terms of both efficiency and accuracy.

Cited by 28SourcecodeScholar
2020

Deep Positional and Relational Feature Learning for Rotation-Invariant Point Cloud Analysis

ECCV 2020poster

In this paper we propose a rotation-invariant deep network for point clouds analysis. Point-based deep networks are commonly designed to recognize roughly aligned 3D shapes based on point coordinates, but suffer from performance drops with shape rotations. Some geometric features, e.g., distances an…

Cited by 44SourcePDFScholar
2020

KLIEP-based Density Ratio Estimation for Semantically Consistent Synthetic to Real Images Adaptation in Urban Traffic Scenes

IROS 2020poster

Synthetic data has been applied in many deep learning based computer vision tasks. Limited performance of algorithms trained solely on synthetic data has been approached with domain adaptation techniques such as the ones based on generative adversarial framework. We demonstrate how adversarial train…

Cited by 1SourceScholar
2020

Learning 3D Semantic Scene Graphs From 3D Indoor Reconstructions

CVPR 2020poster

Scene understanding has been of high interest in computer vision. It encompasses not only identifying objects in a scene, but also their relationships within the given context. With this goal, a recent line of works tackles 3D semantic segmentation and scene layout prediction. In our work we focus o…

Cited by 262PDFScholar
2020

Quaternion Equivariant Capsule Networks for 3D Point Clouds

ECCV 2020poster

We present a 3D capsule module for processing point clouds that is equivariant to 3D rotations and translations, as well as invariant to permutations of the input points. The operator receives a sparse set of local reference frames, computed from an input point cloud and establishes end-to-end trans…

Cited by 112SourcePDFScholar
2020

Self6D: Self-Supervised Monocular 6D Object Pose Estimation

ECCV 2020poster

6D object pose estimation is a fundamental problem in computer vision. Convolutional Neural Networks (CNNs) have recently proven to be capable of predicting reliable 6D pose estimates even from monocular images. Nonetheless, CNNs are identified as being extremely data-driven, and acquiring adequate…

2020

Semantic Image Manipulation Using Scene Graphs

CVPR 2020poster

Image manipulation can be considered a special case of image generation where the image to be produced is a modification of an existing image. Image generation and manipulation have been, for the most part, tasks that operate on raw pixels. However, the remarkable progress in learning rich image and…

Cited by 146PDFcodeScholar
2020

SoftPoolNet: Shape Descriptor for Point Cloud Completion and Classification

ECCV 2020poster

Point clouds are often the default choice for many applications as they exhibit more flexibility and efficiency than volumetric data. Nevertheless, their unorganized nature - points are stored in an unordered way - makes them less suited to be processed by deep learning pipelines. In this paper, we…

Cited by 92SourcePDFScholar
2020

Structure-SLAM: Low-Drift Monocular SLAM in Indoor Environments

RA-L 2020

In this letter a low-drift monocular SLAM method is proposed targeting indoor scenarios, where monocular SLAM often fails due to the lack of textured surfaces. Our approach decouples rotation and translation estimation of the tracking process to reduce the long-term drift in indoor environments. In

Cited by 129SourceScholar
2019

Attention-based Lane Change Prediction

ICRA 2019poster

Lane change prediction of surrounding vehicles is a key building block of path planning. The focus has been on increasing the accuracy of prediction by posing it purely as a function estimation problem at the cost of model understandability. However, the efficacy of any lane change prediction model…

Cited by 60SourceScholar
2019

Crowd-sourced Semantic Edge Mapping for Autonomous Vehicles

IROS 2019poster

Highly accurate maps of the road infrastructure are a crucial cornerstone for self-driving cars to enable navigation in complex traffic scenarios. Traditional methods for creating detailed maps of road environments involve expensive survey vehicles that cannot keep up with the frequent changes in th…

Cited by 34SourceScholar
2019

Explaining the Ambiguity of Object Detection and 6D Pose From Visual Data

ICCV 2019poster

3D object detection and pose estimation from a single image are two inherently ambiguous problems. Oftentimes, objects appear similar from different viewpoints due to shape symmetries, occlusion and repetitive textures. This ambiguity in both detection and pose estimation means that an object instan…

Cited by 138PDFScholar
2019

ForkNet: Multi-Branch Volumetric Semantic Completion From a Single Depth Image

ICCV 2019poster

We propose a novel model for 3D semantic completion from a single depth image, based on a single encoder and three separate generators used to reconstruct different geometric and semantic representations of the original and completed scene, all sharing the same latent space. To transfer information…

Cited by 75PDFScholar
2019

GFrames: Gradient-Based Local Reference Frame for 3D Shape Matching

CVPR 2019oral

We introduce GFrames, a novel local reference frame (LRF) construction for 3D meshes and point clouds. GFrames are based on the computation of the intrinsic gradient of a scalar field defined on top of the input shape. The resulting tangent vector field defines a repeatable tangent direction of the…

Cited by 34PDFScholar
2019

RIO: 3D Object Instance Re-Localization in Changing Indoor Environments

ICCV 2019oral

In this work, we introduce the task of 3D object instance re-localization (RIO): given one or multiple objects in an RGB-D scan, we want to estimate their corresponding 6DoF poses in another 3D scan of the same environment taken at a later point in time. We consider RIO a particularly important task…

Cited by 171PDFcodeScholar
2019

Sampling-Free Epistemic Uncertainty Estimation Using Approximated Variance Propagation

ICCV 2019oral

We present a sampling-free approach for computing the epistemic uncertainty of a neural network. Epistemic uncertainty is an important quantity for the deployment of deep neural networks in safety-critical applications, since it represents how much one can trust predictions on new data. Recently pro…

Cited by 183PDFcodeScholar
2019

Variational Object-Aware 3-D Hand Pose From a Single RGB Image

RA-L 2019

We propose an approach to estimate the 3D pose of a human hand while grasping objects from a single RGB image. Our approach is based on a probabilistic model implemented with deep architectures, which is used for regressing, respectively, the 2D hand joints heat maps and the 3D hand joints coordinat

Cited by 12SourceScholar
2018

BOP: Benchmark for 6D Object Pose Estimation

ECCV 2018poster

We propose a benchmark for 6D pose estimation of a rigid object from a single RGB-D input image. The training data consists of a texture-mapped 3D object model or images of the object in known 6D poses. The benchmark comprises of: i) eight datasets in a unified format that cover different practical…

2018

Distortion-Aware Convolutional Filters for Dense Prediction in Panoramic Images

ECCV 2018poster

There is a high demand of 3D data for 360° panoramic images and videos, pushed by the growing availability on the market of specialized hardware for both capturing (e.g., omnidirectional cameras) as well as visualizing in 3D (e.g., head mounted displays) panoramic images and videos. At the same time…

Cited by 221SourcePDFScholar
2018

Fast and Accurate Semantic Mapping through Geometric-based Incremental Segmentation

IROS 2018poster

We propose an efficient and scalable method for incrementally building a dense, semantically annotated 3D map in real-time. The proposed method assigns class probabilities to each region, not each element (e.g., surfel and voxel), of the 3D map which is built up through a robust SLAM framework and i…

Cited by 51SourceScholar
2018

Fully-Convolutional Point Networks for Large-Scale Point Clouds

ECCV 2018poster

This work proposes a general-purpose, fully-convolutional network architecture for efficiently processing large-scale 3D data. One striking characteristic of our approach is its ability to process unorganized 3D representations such as point clouds as input, then transforming them internally to orde…

2018

Guide Me: Interacting With Deep Networks

CVPR 2018poster

Interaction and collaboration between humans and intelligent machines has become increasingly important as machine learning methods move into real-world applications that involve end users. While much prior work lies at the intersection of natural language and vision, such as image captioning or ima…

Cited by 39SourcePDFScholar
2018

Human Motion Analysis with Deep Metric Learning

ECCV 2018poster

Effectively measuring the similarity between two human motions is necessary for several computer vision tasks such as gait analysis, person identification and action retrieval. Nevertheless, we believe that traditional approaches such as L2 distance or Dynamic Time Warping based on hand-crafted loca…

Cited by 68SourcePDFScholar
2018

Real-Time Fully Incremental Scene Understanding on Mobile Platforms

RA-L 2018

We propose an online RGB-D based scene understanding method for indoor scenes running in real time on mobile devices. First, we incrementally reconstruct the scene via simultaneous localization and mapping and compute a three-dimensional (3-D) geometric segmentation by fusing segments obtained from

Cited by 30SourceScholar
2018

Situation Assessment for Planning Lane Changes: Combining Recurrent Models and Prediction

ICRA 2018poster

One of the greatest challenges towards fully autonomous cars is the understanding of complex and dynamic scenes. Such understanding is needed for planning of maneuvers, especially those that are particularly frequent such as lane changes. While in recent years advanced driver-assistance systems have…

Cited by 41SourceScholar
2017

CNN-SLAM: Real-Time Dense Monocular SLAM With Learned Depth Prediction

CVPR 2017spotlight

Given the recent advances in depth prediction from Convolutional Neural Networks (CNNs), this paper investigates how predicted depth maps from a deep neural network can be deployed for the goal of accurate and dense monocular reconstruction. We propose a method where CNN-predicted dense depth maps a…

Cited by 1024PDFcodeScholar
2017

Learning in an Uncertain World: Representing Ambiguity Through Multiple Hypotheses

ICCV 2017poster

Many prediction tasks contain uncertainty. In some cases, uncertainty is inherent in the task itself. In future prediction, for example, many distinct outcomes are equally valid. In other cases, uncertainty arises from the way data is labeled. For example, in object detection, many objects of intere…

Cited by 235PDFScholar
2017

Long Short-Term Memory Kalman Filters: Recurrent Neural Estimators for Pose Regularization

ICCV 2017poster

One-shot pose estimation for tasks such as body joint localization, camera pose estimation, and object tracking are generally noisy, and temporal filters have been extensively used for regularization. One of the most widely-used methods is the Kalman filter, which is both extremely simple and genera…

Cited by 237PDFScholar
2017

Real-Time 3D Model Tracking in Color and Depth on a Single CPU Core

CVPR 2017poster

We present a novel method to track 3D models in color and depth data. To this end, we introduce approximations that accelerate the state-of-the-art in region-based tracking by an order of magnitude while retaining similar accuracy. Furthermore, we show how the method can be made more robust in the p…

Cited by 51PDFScholar
2017

SSD-6D: Making RGB-Based 3D Detection and 6D Pose Estimation Great Again

ICCV 2017oral

We present a novel method for detecting 3D model instances and estimating their 6D poses from RGB data in a single shot. To this end, we extend the popular SSD paradigm to cover the full 6D pose space and train on synthetic model data only. Our approach competes or surpasses current state-of-the-art…

Cited by 1284PDFScholar
2016

Incremental scene understanding on dense SLAM

IROS 2016poster

We present an architecture for online, incremental scene modeling which combines a SLAM-based scene understanding framework with semantic segmentation and object pose estimation. The core of this approach comprises a probabilistic inference scheme that predicts semantic labels for object hypotheses…

Cited by 36SourceScholar
2016

Sensor substitution for video-based action recognition

IROS 2016poster

There are many applications where domain-specific sensing, such as accelerometers, kinematics, or force sensing, provide unique and important information for control or for analysis of motion. However, it is not always the case that these sensors can be deployed or accessed beyond laboratory environ…

Cited by 34SourceScholar
2016

When 2.5D is not enough: Simultaneous reconstruction, segmentation and recognition on dense SLAM

ICRA 2016

While the main trend of 3D object recognition has been to infer object detection from single views of the scene - i.e., 2.5D data - this work explores the direction on performing object recognition on 3D data that is reconstructed from multiple viewpoints, under the conjecture that such data can imp

Cited by 95SourceScholar
2015

A Versatile Learning-Based 3D Temporal Tracker: Scalable, Robust, Online

ICCV 2015poster

This paper proposes a temporal tracking algorithm based on Random Forest that uses depth images to estimate and track the 3D pose of a rigid object in real-time. Compared to the state of the art aimed at the same goal, our algorithm holds important attributes such as high robustness against holes an…

Cited by 87PDFScholar
2015

Learning a Descriptor-Specific 3D Keypoint Detector

ICCV 2015poster

Keypoint detection represents the first stage in the majority of modern computer vision pipelines based on automatically established correspondences between local descriptors. However, no standard solution has emerged yet in the case of 3D data such as point clouds or meshes, which exhibit high vari…

Cited by 54PDFScholar