← Search

Nassir Navab

168 accepted papers

2026

Conformable Convolution for Topologically Constrained Learning of Complex Anatomical Structures

AAAI 2026technical

While conventional computer vision emphasizes pixel-level and feature-based objectives, medical image analysis of intricate biological structures necessitates explicit representation of their complex topological properties. Despite their successes, deep learning models often struggle to accurately c

Cited by 0SourcePDFScholar
2026

DisPOSE: Projected Polystochastic Diffusion for Self-Supervised Multi-View 3D Human Pose Estimation

ICML 2026poster

Recovering 3D human poses for multiple individuals from different camera views is a fundamental bottleneck for analyzing interacting behaviors. Existing self-supervised approaches leverage synthetic catalogues of 3D poses; however, this leads to poor generalization in real-world scenarios due to dis…

Cited by 0SourceScholar
2026

Foundation Visual Encoders Are Secretly Few-Shot Anomaly Detectors

ICLR 2026poster

Few-shot anomaly detection streamlines and simplifies industrial safety inspection. However, limited samples make accurate differentiation between normal and abnormal features challenging, and even more so under category-agnostic conditions. Large-scale pre-training of foundation visual encoders has…

Cited by 0SourcecodeScholar
2026

From Panel to Pixel: Zoom-In Vision-Language Pretraining from Biomedical Scientific Literature

CVPR 2026

There is growing interest in biomedical vision--language models trained on scientific literature. However, most pipelines compress rich multi-panel figures and long captions into coarse figure-level pairs, discarding the fine-grained correspondences clinicians rely on when zooming into local structu

Cited by 0SourceScholar
2026

Future Dynamic 3D Reconstruction: A 3D World Model with Disentangled Ego-Motion

ICML 2026poster

Forecasting the evolution of dynamic environments is crucial for autonomous agents. While generative world models have recently achieved high photorealism in 2D video synthesis, by mixing within the image plane ego-motion and environmental dynamics, they exhibit physical inconsistencies, such as mor…

Cited by 0SourceScholar
2026

Gaze-Guided Robotic Vascular Ultrasound Leveraging Human Intention Estimation

ICRA 2026poster

Medical ultrasound (US) has been widely used to examine vascular structure in modern clinical practice. However, the traditional US examination often faces challenges related to inter- and intra-operator variation. The robotic ultrasound system (RUSS) appears as a potential solution for such challen…

2026

Language Agents for Hypothesis-driven Clinical Decision Making with Reinforcement Learning

ICLR 2026poster

Clinical decision-making is a dynamic, interactive, and cyclic process where doctors have to repeatedly decide on which clinical action to perform and consider newly uncovered information for diagnosis and treatment. Large Language Models (LLMs) have the potential to support clinicians in this proce…

Cited by 0SourcecodeScholar
2026

RAG-RUSS: A Retrieval-Augmented Robotic Ultrasound for Autonomous Carotid Examination

ICRA 2026poster

Robotic ultrasound (US) has recently attracted considerable attention as a means to overcome the limitations of conventional US examinations, such as the strong operator dependence. However, the decision-making process of existing methods is often either rule-based or relies on end-to-end learning m…

2026

Rewarding Doubt: A Reinforcement Learning Approach to Calibrated Confidence Expression of Large Language Models

ICLR 2026poster

A safe and trustworthy use of Large Language Models (LLMs) requires an accurate expression of confidence in their answers. We propose a novel Reinforcement Learning approach that allows to directly fine-tune LLMs to express calibrated confidence estimates alongside their answers to factual questions…

Cited by 0SourcecodeScholar
2026

SPEECHCT-CLIP: DISTILLING TEXT-IMAGE KNOWLEDGE TO SPEECH FOR VOICE-NATIVE MULTIMODAL CT ANALYSIS

ICASSP 2026oral

Spoken communication plays a central role in clinical workflows. In radiology, for example, most reports are created through dictation. Yet, nearly all medical AI systems rely exclusively on written text. In this work, we address this gap by exploring the feasibility of learning visual-language repr…

Cited by 0SourcePDFScholar
2026

UnReflectAnything: RGB-Only Highlight Removal by Rendering Synthetic Specular Supervision

CVPR 2026

Specular highlights distort appearance, obscure texture, and hinder geometric reasoning in both natural and surgical imagery. We present UnReflectAnything, an RGB-only framework that removes highlights from a single image by predicting a highlight map together with a reflection-free diffuse reconstr

Cited by 1SourcecodeScholar
2026

Where It Moves, It Matters: Referring Surgical Instrument Segmentation via Motion

AAAI 2026technical

Enabling intuitive, language-driven interaction with surgical scenes is a critical step toward intelligent operating rooms and autonomous surgical robotic assistance. However, the task of referring segmentation, localizing surgical instruments based on natural language descriptions, remains underexp

Cited by 0SourcePDFScholar
2025

Design and Effectiveness of Virtual Monitors and AR-Based Endoscope Control for Robotically Assisted Laparoscopic Surgery

ICRA 2025

Managing indirect access in laparoscopy as a minimally invasive procedure poses challenges to physicians. In particular, an endoscope must be navigated to achieve adequate visualization of the surgical anatomy, while coping with unergonomic poses, tremor, and fatigue. Furthermore, the alignment of v

Cited by 1SourceScholar
2025

DynaMoN: Motion-Aware Fast and Robust Camera Localization for Dynamic Neural Radiance Fields

RA-L 2025

The accurate reconstruction of dynamic scenes with neural radiance fields is significantly dependent on the estimation of camera poses. Widely used structure-from-motion pipelines encounter difficulties in accurately tracking the camera trajectory when faced with separate dynamics of the scene conte

Cited by 21SourcecodeScholar
2025

ESCAPE: Equivariant Shape Completion via Anchor Point Encoding

CVPR 2025poster

Shape completion, a crucial task in 3D computer vision, involves predicting and filling the missing regions of scanned or partially observed objects. Current methods expect known pose or canonical coordinates and do not perform well under varying rotations, limiting their real-world applicability. W…

2025

EgoExOR: An Ego-Exo-Centric Operating Room Dataset for Surgical Activity Understanding

NeurIPS 2025poster

Operating rooms (ORs) demand precise coordination among surgeons, nurses, and equipment in a fast-paced, occlusion-heavy environment, necessitating advanced perception models to enhance safety and efficiency. Existing datasets either provide partial egocentric views or sparse exocentric multi-view c…

Cited by 0SourcecodeScholar
2025

FB-Diff: Fourier Basis-guided Diffusion for Temporal Interpolation of 4D Medical Imaging

ICCV 2025poster

The temporal interpolation task for 4D medical imaging, plays a crucial role in clinical practice of respiratory motion modeling. Following the simplified linear-motion hypothesis, existing approaches adopt optical flow-based models to interpolate intermediate frames. However, realistic respiratory…

2025

Forecasting Continuous Non-Conservative Dynamical Systems in SO(3)

ICCV 2025poster

Modeling the rotation of moving objects is a fundamental task in computer vision, yet SO(3) extrapolation still presents numerous challenges: (1) unknown quantities such as the moment of inertia complicate dynamics, (2) the presence of external forces and torques can lead to non-conservative kinemat…

Cited by 0SourcePDFScholar
2025

GCE-Pose: Global Context Enhancement for Category-level Object Pose Estimation

CVPR 2025poster

A key challenge in model-free category-level pose estimation is the extraction of contextual object features that generalize across varying instances within a specific category. Recent approaches leverage foundational features to capture semantic and geometry cues from data. However, these approache…

Cited by 2SourcePDFScholar
2025

Improving Probe Localization for Freehand 3D Ultrasound Using Lightweight Cameras

ICRA 2025

Ultrasound (US) probe localization relative to the examined subject is essential for freehand 3D US imaging, which offers significant clinical value due to its affordability and unrestricted field of view. However, existing methods often rely on expensive tracking systems or bulky probes, while rece

Cited by 1SourcecodeScholar
2025

Intraoperative Trocar-Based Eyeball Rotation Estimation Using Only 2D Microscope Images

ICRA 2025

In ophthalmic surgery, surgeons or robots manipulate a light probe and an instrument around two separated trocars following sclerotomy to achieve orbital control for eyeball pose adjustment and subsequent surgical tasks referring to microscope frames. However, current methods face significant challe

Cited by 0SourceScholar
2025

Latent Drifting in Diffusion Models for Counterfactual Medical Image Synthesis

CVPR 2025highlight

Scaling by training on large datasets has been shown to enhance the quality and fidelity of image generation and manipulation with diffusion models; however, such large datasets are not always accessible in medical imaging due to cost and privacy issues, which contradicts one of the main application…

Cited by 0SourcePDFScholar
2025

MM-OR: A Large Multimodal Operating Room Dataset for Semantic Understanding of High-Intensity Surgical Environments

CVPR 2025poster

Operating rooms (ORs) are complex, high-stakes environments requiring precise understanding of interactions among medical staff, tools, and equipment for enhancing surgical assistance, situational awareness, and patient safety. Current datasets fall short in scale, realism and do not capture the mul…

2025

Medical Multimodal Model Stealing Attacks via Adversarial Domain Alignment

AAAI 2025technical

Medical multimodal large language models (MLLMs) are becoming an instrumental part of healthcare systems, assisting medical personnel with decision making and results analysis. Models for radiology report generation are able to interpret medical imagery, thus reducing the workload of radiologists. A…

Cited by 1SourcePDFScholar
2025

OphCLIP: Hierarchical Retrieval-Augmented Learning for Ophthalmic Surgical Video-Language Pretraining

ICCV 2025poster

Vision-language pretraining (VLP) enables open-world generalization beyond predefined labels, a critical capability in surgery due to the diversity of procedures, instruments, and patient anatomies. However, applying VLP to ophthalmic surgery presents unique challenges, including limited vision-lang…

2025

Pre-Surgical Planner for Robot-Assisted Vitreoretinal Surgery: Integrating Eye Posture, Robot Position and Insertion Point

ICRA 2025

Several robotic frameworks have been recently developed to assist ophthalmic surgeons in performing complex vitreoretinal procedures such as subretinal injection of advanced therapeutics. These surgical robots show promising capabilities; however, most of them have to limit their working volume to a

Cited by 0SourceScholar
2025

RayPose: Ray Bundling Diffusion for Template Views in Unseen 6D Object Pose Estimation

ICCV 2025poster

Typical template-based object pose pipelines estimate the pose by retrieving the closest matching template and aligning it with the observed image. However, failure to retrieve the correct template often leads to inaccurate pose predictions. To address this, we reformulate template-based object pose…

Cited by 0SourcePDFScholar
2025

Real-Time Deformation-Aware Control for Autonomous Robotic Subretinal Injection Under iOCT Guidance

ICRA 2025

Robotic platforms provide consistent and precise tool positioning that significantly enhances retinal microsurgery. Integrating such systems with intraoperative optical coherence tomography (iOCT) enables image-guided robotic interventions, allowing autonomous performance of advanced treatments, suc

Cited by 10SourceScholar
2025

Shape Completion and Real-Time Visualization in Robotic Ultrasound Spine Acquisitions

IROS 2025

Ultrasound (US) imaging is increasingly used in spinal procedures due to its real-time, radiation-free capabilities; however, its effectiveness is hindered by shadowing artifacts that obscure deeper tissue structures. Traditional approaches, such as CT-to-US registration, incorporate anatomical info

Cited by 4SourceScholar
2025

Tactile-Guided Robotic Ultrasound: Mapping Preplanned Scan Paths for Intercostal Imaging

IROS 2025

Medical ultrasound (US) imaging is widely used in clinical examinations due to its portability, real-time capability, and radiation-free nature. To address inter- and intra-operator variability, robotic ultrasound systems have gained increasing attention. However, their application in challenging in

Cited by 1SourceScholar
2025

Video-Rate 4D OCT Segmentation Based on Motion-Aware Probabilistic A-Scan Sampling

IROS 2025

Recent advancements in robotic eye surgery and intraoperative 4D optical coherence tomography (iOCT) imaging could enable fully or partially autonomous robotic procedures and enhanced surgical visualization. A fundamental requirement for such applications is rapid semantic segmentation of intraopera

Cited by 0SourceScholar
2025

VoxNeRF: Bridging Voxel Representation and Neural Radiance Fields for Enhanced Indoor View Synthesis

RA-L 2025

The generation of high-fidelity view synthesis is essential for robotic navigation and interaction but remains challenging, particularly in indoor environments and real-time scenarios. Existing techniques often require significant computational resources for both training and rendering, and they fre

Cited by 2SourceScholar
2024

AiAReSeg: Catheter Detection and Segmentation in Interventional Ultrasound using Transformers

ICRA 2024poster

This work proposes a state-of-the-art transformer architecture to detect and segment catheters in axial interventional Ultrasound image sequences. The network architecture was inspired by the Attention in Attention mechanism, temporal tracking networks, and introduced a novel 3D segmentation head th…

Cited by 4SourceScholar
2024

Analyzing Accessibility in Robot-Assisted Vitreoretinal Surgery: Integrating Eye Posture and Robot Position

ICRA 2024poster

Several robotic frameworks have been recently developed to assist ophthalmic surgeons in performing complex vitreoretinal procedures such as subretinal injection. However, in order to intuitively integrate robots into the surgical workflow, it is crucial to emphasize that an accessibility analysis f…

Cited by 2SourceScholar
2024

CathFlow: Self-Supervised Segmentation of Catheters in Interventional Ultrasound Using Optical Flow and Transformers

IROS 2024poster

In minimally invasive endovascular procedures, contrast-enhanced angiography remains the most robust imaging technique, but exposes patients and surgeons to prolonged radiation. Alternatives such as ultrasound are difficult to interpret, are highly prone to artifacts and noise, and vary in quality,…

Cited by 2SourceScholar
2024

Colibri5: Real-Time Monocular 5-DoF Trocar Pose Tracking for Robot-Assisted Vitreoretinal Surgery

ICRA 2024poster

Retinal surgery is a complex medical procedure that requires high precision dexterity to perform delicate instrument maneuvers with sub-millimeter accuracy. Minimizing the manual tremor and achieving precise and repeatable execution of surgical tasks has motivated the development of robotic platform…

Cited by 2SourceScholar
2024

EchoScene: Indoor Scene Generation via Information Echo over Scene Graph Diffusion

ECCV 2024poster

"We present EchoScene, an interactive and controllable generative model that generates 3D indoor scenes on scene graphs. EchoScene leverages a dual-branch diffusion model that dynamically adapts to scene graphs. Existing methods struggle to handle scene graphs due to varying numbers of nodes, multip…

Cited by 23SourcePDFScholar
2024

Envibroscope: Real-Time Monitoring and Prediction of Environmental Motion for Enhancing Safety in Robot-Assisted Microsurgery

ICRA 2024

Several robotic systems have emerged in the recent past to enhance the precision of micro-surgeries such as retinal procedures. Significant advancements have recently been achieved to increase the precision of such systems beyond surgeon capabilities. However, little attention has been paid to the i

Cited by 2SourceScholar
2024

Exploring the Needle Tip Interaction Force with Retinal Tissue Deformation in Vitreoretinal Surgery

ICRA 2024poster

Recent advancements in age-related macular degeneration treatments necessitate precision delivery into the subretinal space, emphasizing minimally invasive procedures targeting the retinal pigment epithelium (RPE)-Bruch’s membrane complex without causing trauma. Even for skilled surgeons, the inhere…

Cited by 3SourceScholar
2024

EyeLS: Shadow-Guided Instrument Landing System for Target Approaching in Robotic Eye Surgery

RA-L 2024

Robotic ophthalmic surgery is an emerging technology to facilitate high-precision interventions such as subretinal injection and removing swinging tissues in retinal detachment using microscopy and iOCT. However, locating the instrument tip outside iOCT's range-limited ROI is challenging, especially

Cited by 3SourceScholar
2024

HouseCat6D - A Large-Scale Multi-Modal Category Level 6D Object Perception Dataset with Household Objects in Realistic Scenarios

CVPR 2024highlight

Estimating 6D object poses is a major challenge in 3D computer vision. Building on successful instance-level approaches research is shifting towards category-level pose estimation for practical applications. Current category-level datasets however fall short in annotation quality and pose variety. A…

2024

Hybrid Functional Maps for Crease-Aware Non-Isometric Shape Matching

CVPR 2024poster

Non-isometric shape correspondence remains a fundamental challenge in computer vision. Traditional methods using Laplace-Beltrami operator (LBO) eigenmodes face limitations in characterizing high-frequency extrinsic shape changes like bending and creases. We propose a novel approach of combining the…

Cited by 5SourcePDFScholar
2024

Implicit Neural Representations for Breathing-compensated Volume Reconstruction in Robotic Ultrasound

ICRA 2024poster

Ultrasound (US) imaging is widely used in diagnosing and staging abdominal diseases due to its lack of non-ionizing radiation and prevalent availability. However, significant inter-operator variability and inconsistent image acquisition hinder the widespread adoption of extensive screening programs.…

Cited by 4SourceScholar
2024

Improving Self-Supervised Learning of Transparent Category Poses With Language Guidance and Implicit Physical Constraints

RA-L 2024

Accurate object pose estimation is crucial for robotic applications and recent trends in category-level pose estimation show great potential for applications encountering a large variety of similar objects, often encountered in home environments. While common in such environments, photometrically ch

Cited by 1SourceScholar
2024

Intraocular Reflection Modeling and Avoidance Planning in Image-Guided Ophthalmic Surgeries

IROS 2024

Intuitive enhancement of surgical precision in robotic retinal surgery highly depends on the stable acquisition of intraocular imaging data. Such acquisition requires segmenting intraocular components, especially instrument-tip positions, to achieve state estimation and subsequent navigation and mot

Cited by 0SourceScholar
2024

MatchU: Matching Unseen Objects for 6D Pose Estimation from RGB-D Images

CVPR 2024poster

Recent learning methods for object pose estimation require resource-intensive training for each individual object instance or category hampering their scalability in real applications when confronted with previously unseen objects. In this paper we propose MatchU a Fuse-Describe-Match strategy for 6…

Cited by 10SourcePDFScholar
2024

Physics-Encoded Graph Neural Networks for Deformation Prediction under Contact

ICRA 2024poster

In robotics, it’s crucial to understand object deformation during tactile interactions. A precise understanding of deformation can elevate robotic simulations and have broad implications across different industries. We introduce a method using Physics-Encoded Graph Neural Networks (GNNs) for such pr…

Cited by 4SourceScholar
2024

Procedure-Aware Surgical Video-language Pretraining with Hierarchical Knowledge Augmentation

NeurIPS 2024spotlight

Surgical video-language pretraining (VLP) faces unique challenges due to the knowledge domain gap and the scarcity of multi-modal data. This study aims to bridge the gap by addressing issues regarding textual information loss in surgical lecture videos and the spatial-temporal challenges of surgical…

2024

RIDE: Self-Supervised Learning of Rotation-Equivariant Keypoint Detection and Invariant Description for Endoscopy

ICRA 2024poster

Unlike in natural images, in endoscopy there is no clear notion of an up-right camera orientation. Endoscopic videos therefore often contain large rotational motions, which require keypoint detection and description algorithms to be robust to these conditions. While most classical methods achieve ro…

Cited by 6SourceScholar
2024

SCRREAM : SCan, Register, REnder And Map: A Framework for Annotating Accurate and Dense 3D Indoor Scenes with a Benchmark

NeurIPS 2024poster

Traditionally, 3d indoor datasets have generally prioritized scale over ground-truth accuracy in order to obtain improved generalization. However, using these datasets to evaluate dense geometry tasks, such as depth rendering, can be problematic as the meshes of the dataset are often incomplete and…

2024

SG-Bot: Object Rearrangement via Coarse-to-Fine Robotic Imagination on Scene Graphs

ICRA 2024poster

Object rearrangement is pivotal in robotic-environment interactions, representing a significant capability in embodied AI. In this paper, we present SG-Bot, a novel rearrangement framework that utilizes a coarse-to-fine scheme with a scene graph as the scene representation. Unlike previous methods t…

Cited by 25SourceScholar
2024

SecondPose: SE(3)-Consistent Dual-Stream Feature Fusion for Category-Level Pose Estimation

CVPR 2024poster

Category-level object pose estimation aiming to predict the 6D pose and 3D size of objects from known categories typically struggles with large intra-class shape variation. Existing works utilizing mean shapes often fall short of capturing this variation. To address this issue we present SecondPose…

2024

Shadow Maintenance for Automatic Light-Probe Control in Ophthalmic Surgeries Using Only 2D information

IROS 2024poster

In ophthalmic surgeries, the light probe is responsible for providing safe intraocular illumination and ensuring the visibility of the instrument and its shadow as the only available reference for qualitative depth estimation and landing point prediction in fundus microscopic images. To achieve sust…

Cited by 0SourceScholar
2024

Shadow-Based 3D Pose Estimation of Intraocular Instrument Using Only 2D Images

ICRA 2024poster

In ophthalmic surgeries, such as vitreoretinal operations, surgeons rely on imaging systems, primarily microscopes, for real-time instrument monitoring and motion planning. However, novice surgeons struggle to extract 3D instrument positions from 2D microscope frames, necessitating extensive trial-a…

Cited by 0SourceScholar
2024

Uncertainty-Aware Contextual Visualization for Human Supervision of OCT-Guided Autonomous Robotic Subretinal Injection

ICRA 2024poster

The injection of therapeutic agents into the sub-retinal space might allow improved treatment of age-related macular degeneration. Various robotic systems have been developed to achieve the required precision and, in combination with intraoperative Optical Coherence Tomography (iOCT) imaging, method…

Cited by 2SourceScholar
2023

CommonScenes: Generating Commonsense 3D Indoor Scenes with Scene Graph Diffusion

NeurIPS 2023poster

Controllable scene synthesis aims to create interactive environments for numerous industrial use cases. Scene graphs provide a highly suitable interface to facilitate these applications by abstracting the scene context in a compact manner. Existing methods, reliant on retrieval from extensive databa…

2023

Dynamic Hyperbolic Attention Network for Fine Hand-object Reconstruction

ICCV 2023poster

Reconstructing both objects and hands in 3D from a single RGB image is complex. Existing methods rely on manually defined hand-object constraints in Euclidean space, leading to suboptimal feature learning. Compared with Euclidean space, hyperbolic space better preserves the geometric properties of m…

Cited by 14PDFScholar
2023

Feature-Based Electromagnetic Tracking Registration Using Bioelectric Sensing

RA-L 2023

Catheter tracking is essential during minimally invasive endovascular procedures, and Electromagnetic (EM) tracking is a widely used technology for this purpose. When preoperative patient images are available, they can be used to guide EM-tracked interventions. However, a registration step between p

Cited by 3SourceScholar
2023

IPCC-TP: Utilizing Incremental Pearson Correlation Coefficient for Joint Multi-Agent Trajectory Prediction

CVPR 2023poster

Reliable multi-agent trajectory prediction is crucial for the safe planning and control of autonomous systems. Compared with single-agent cases, the major challenge in simultaneously processing multiple agents lies in modeling complex social interactions caused by various driving intentions and road…

Cited by 19SourcePDFScholar
2023

Incremental 3D Semantic Scene Graph Prediction From RGB Sequences

CVPR 2023poster

3D semantic scene graphs are a powerful holistic representation as they describe the individual objects and depict the relation between them. They are compact high-level graphs that enable many tasks requiring scene reasoning. In real-world settings, existing 3D estimation methods produce robust pre…

2023

MonoGraspNet: 6-DoF Grasping with a Single RGB Image

ICRA 2023poster

6-DoF robotic grasping is a long-lasting but un-solved problem. Recent methods utilize strong 3D networks to extract geometric grasping representations from depth sensors, demonstrating superior accuracy on common objects but performing unsatisfactorily on photometrically challenging objects, e.g.,…

Cited by 39SourceScholar
2023

Motion Magnification in Robotic Sonography: Enabling Pulsation-Aware Artery Segmentation

IROS 2023poster

Ultrasound (US) imaging is widely used for diagnosing and monitoring arterial diseases, mainly due to the advantages of being non-invasive, radiation-free, and real-time. In order to provide additional information to assist clinicians in diagnosis, the tubular structures are often segmented from US…

Cited by 12SourcecodeScholar
2023

On the Importance of Accurate Geometry Data for Dense 3D Vision Tasks

CVPR 2023poster

Learning-based methods to solve dense 3D vision problems typically train on 3D sensor data. The respectively used principle of measuring distances provides advantages and drawbacks. These are typically not compared nor discussed in the literature due to a lack of multi-modal datasets. Texture-less r…

2023

Robotic Navigation Autonomy for Subretinal Injection via Intelligent Real-Time Virtual iOCT Volume Slicing

ICRA 2023poster

In the last decade, various robotic platforms have been introduced that could support delicate retinal surgeries. Concurrently, to provide semantic understanding of the surgical area, recent advances have enabled microscope-integrated intraoperative Optical Coherent Tomography (iOCT) with high-resol…

Cited by 24SourceScholar
2023

Robust Monocular Depth Estimation under Challenging Conditions

ICCV 2023accepted

While state-of-the-art monocular depth estimation approaches achieve impressive results in ideal settings, they are highly unreliable under challenging illumination and weather conditions, such as at nighttime or in the presence of rain. In this paper, we uncover these safety-critical issues and tac…

2023

Segmenting Known Objects and Unseen Unknowns without Prior Knowledge

ICCV 2023poster

Panoptic segmentation methods assign a known class to each pixel given in input. Even for state-of-the-art approaches, this inevitably enforces decisions that systematically lead to wrong predictions for objects outside the training categories. However, robustness against out-of-distribution samples…

Cited by 10PDFcodeScholar
2023

Skeleton Graph-Based Ultrasound-CT Non-Rigid Registration

RA-L 2023

Autonomous ultrasound (US) scanning has attracted increased attention, and it has been seen as a potential solution to overcome the limitations of conventional US examinations, such as inter-operator variations. However, it is still challenging to autonomously and accurately transfer a planned scan

Cited by 14SourceScholar
2023

SupeRGB-D: Zero-Shot Instance Segmentation in Cluttered Indoor Environments

RA-L 2023

Object instance segmentation is a key challenge for indoor robots navigating cluttered environments with many small objects. Limitations in 3D sensing capabilities often make it difficult to detect every possible object. While deep learning approaches may be effective for this problem, manually anno

Cited by 14SourcecodeScholar
2023

TexPose: Neural Texture Learning for Self-Supervised 6D Object Pose Estimation

CVPR 2023poster

In this paper, we introduce neural texture learning for 6D object pose estimation from synthetic data and a few unlabelled real images. Our major contribution is a novel learning scheme which removes the drawbacks of previous works, namely the strong dependency on co-modalities or additional refinem…

Cited by 40SourcePDFScholar
2023

Thoracic Cartilage Ultrasound-CT Registration Using Dense Skeleton Graph

IROS 2023poster

Autonomous ultrasound (US) imaging has gained increased interest recently, and it has been seen as a potential solution to overcome the limitations of free-hand US exami-nations, such as inter-operator variations. However, it is still challenging to accurately map planned paths from a generic atlas…

Cited by 7SourcecodeScholar
2023

Towards Long-Term Retrieval-Based Visual Localization in Indoor Environments With Changes

RA-L 2023

Visual localization is a challenging task due to the presence of illumination changes, occlusion, and perception from novel viewpoints. Re-localizing the camera pose in long-term setups raises difficulties caused by changes in scene appearance and geometry introduced by human or natural deterioratio

Cited by 12SourceScholar
2022

3D-VField: Adversarial Augmentation of Point Clouds for Domain Generalization in 3D Object Detection

CVPR 2022poster

As 3D object detection on point clouds relies on the geometrical relationships between the points, non-standard object shapes can hinder a method's detection capability. However, in safety-critical settings, robustness to out-of-domain and long-tail samples is fundamental to circumvent dangerous iss…

Cited by 69PDFcodeScholar
2022

A Variational Bayesian Method for Similarity Learning in Non-Rigid Image Registration

CVPR 2022poster

We propose a novel variational Bayesian formulation for diffeomorphic non-rigid registration of medical images, which learns in an unsupervised way a data-specific similarity metric. The proposed framework is general and may be used together with many existing image registration models. We evaluate…

Cited by 12PDFcodeScholar
2022

Acoustic Shadowing Aware Robotic Ultrasound: Lighting up the Dark

RA-L 2022

Ultrasound imaging is becoming more prevalent in clinical practice and research. To counteract the drawbacks of high user-dependency and difficult interpretability, ultrasound probes can be attached to robotic arms, enabling an increase in accuracy and repeatability. Currently, robotic ultrasound sc

Cited by 9SourceScholar
2022

Bending Graphs: Hierarchical Shape Matching Using Gated Optimal Transport

CVPR 2022poster

Shape matching has been a long-studied problem for the computer graphics and vision community. The objective is to predict a dense correspondence between meshes that have a certain degree of deformation. Existing methods either consider the local description of sampled points or discover corresponde…

Cited by 24PDFcodeScholar
2022

CertainNet: Sampling-Free Uncertainty Estimation for Object Detection

RA-L 2022

Estimating the uncertainty of a neural network plays a fundamental role in safety-critical settings. In perception for autonomous driving, measuring the uncertainty means providing additional calibrated information to downstream tasks, such as path planning, that can use it towards safe navigation.

Cited by 29SourceScholar
2022

CloudAttention: Efficient Multi-Scale Attention Scheme For 3D Point Cloud Learning

IROS 2022poster

Processing 3D data efficiently has always been a challenge. Spatial operations on large-scale point clouds, stored as sparse data, require extra cost. Attracted by the success of transformers, researchers are using multi-head attention for vision tasks. However, attention calculations in transformer…

Cited by 5SourcecodeScholar
2022

ColibriDoc: an Eye-in-Hand Autonomous Trocar Docking System

ICRA 2022poster

Retinal surgery is a complex medical procedure that requires exceptional expertise and dexterity. For this purpose, several robotic platforms are currently under development to enable or improve the outcome of microsurgical tasks. Since the control of such robots is often designed for navigation ins…

Cited by 18SourceScholar
2022

DA${2}$ Dataset: Toward Dexterity-Aware Dual-Arm Grasping

RA-L 2022

In this paper, we introduce DA <inline-formula xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink"><tex-math notation="LaTeX">$^{2}$</tex-math></inline-formula> , the first large-scale dual-arm dexterity-aware dataset for the generation of optimal bimanual grasp

Cited by 21SourceScholar
2022

GPV-Pose: Category-Level Object Pose Estimation via Geometry-Guided Point-Wise Voting

CVPR 2022poster

While 6D object pose estimation has recently made a huge leap forward, most methods can still only handle a single or a handful of different objects, which limits their applications. To circumvent this problem, category-level object pose estimation has recently been revamped, which aims at predictin…

Cited by 147PDFcodeScholar
2022

Object-Aware Monocular Depth Prediction With Instance Convolutions

RA-L 2022

With the advent of deep learning, estimating depth from a single RGB image has recently received a lot of attention, being capable of empowering many different applications ranging from path planning for robotics to computational cinematography. Nevertheless,while the depth maps are in their entiret

Cited by 3SourcecodeScholar
2022

PhoCaL: A Multi-Modal Dataset for Category-Level Object Pose Estimation With Photometrically Challenging Objects

CVPR 2022poster

Object pose estimation is crucial for robotic applications and augmented reality. Beyond instance level 6D object pose estimation methods, estimating category-level pose and shape has become a promising trend. As such, a new research field needs to be supported by well-designed datasets. To provide…

Cited by 54PDFScholar
2022

RSV: Robotic Sonography for Thyroid Volumetry

RA-L 2022

In nuclear medicine, radioiodine therapy is prescribed to treat diseases like hyperthyroidism. The calculation of the prescribed dose depends, amongst other factors, on the thyroid volume. This is currently estimated using conventional 2D ultrasound imaging. However, this modality is inherently user

Cited by 24SourceScholar
2022

Towards Autonomous Atlas-Based Ultrasound Acquisitions in Presence of Articulated Motion

RA-L 2022

Robotic ultrasound (US) imaging aims at overcoming some of the limitations of free-hand US examinations, e.g. difficulty in guaranteeing intra- and inter-operator repeatability. However, due to anatomical and physiological variations between patients and relative movement of anatomical substructures

Cited by 36SourcecodeScholar
2022

VesNet-RL: Simulation-Based Reinforcement Learning for Real-World US Probe Navigation

RA-L 2022

Ultrasound (US) is one of the most common medical imaging modalities since it is radiation-free, low-cost, and real-time. In freehand US examinations, sonographers often navigate a US probe to visualize standard examination planes with rich diagnostic information. However, reproducibility and stabil

Cited by 53SourcecodeScholar
2022

ZebraPose: Coarse To Fine Surface Encoding for 6DoF Object Pose Estimation

CVPR 2022poster

Establishing correspondences from image to 3D has been a key task of 6DoF object pose estimation for a long time. To predict pose more accurately, deeply learned dense maps replaced sparse templates. Dense methods also improved pose estimation in the presence of occlusion. More recently researchers…

Cited by 176PDFcodeScholar
2021

DemoGrasp: Few-Shot Learning for Robotic Grasping with Human Demonstration

IROS 2021poster

The ability to successfully grasp objects is crucial in robotics, as it enables several interactive downstream applications. To this end, most approaches either compute the full 6D pose for the object of interest or learn to predict a set of grasping points. While the former approaches do not scale…

Cited by 38SourceScholar
2021

Fine-Grained Neural Network Explanation by Identifying Input Features with Predictive Information

NeurIPS 2021poster

One principal approach for illuminating a black-box neural network is feature attribution, i.e. identifying the importance of input features for the network’s prediction. The predictive information of features is recently proposed as a proxy for the measure of their importance. So far, the predictiv…

2021

Graph-to-3D: End-to-End Generation and Manipulation of 3D Scenes Using Scene Graphs

ICCV 2021poster

Controllable scene synthesis consists of generating 3D information that satisfy underlying specifications. Thereby, these specifications should be abstract, i.e. allowing easy user interaction, whilst providing enough interface for detailed control. Scene graphs are representations of a scene, compo…

Cited by 85PDFcodeScholar
2021

Motion-Aware Robotic 3D Ultrasound

ICRA 2021poster

Robotic three-dimensional (3D) ultrasound (US) imaging has been employed to overcome the drawbacks of traditional US examinations, such as high inter-operator variability and lack of repeatability. However, object movement remains a challenge as unexpected motion decreases the quality of the 3D comp…

Cited by 29SourceScholar
2021

Neural Response Interpretation Through the Lens of Critical Pathways

CVPR 2021poster

Is critical input information encoded in specific sparse pathways within the neural network? In this work, we discuss the problem of identifying these critical pathways and subsequently leverage them for interpreting the network's response to an input. The pruning objective --- selecting the smalles…

Cited by 41PDFcodeScholar
2021

Panoster: End-to-End Panoptic Segmentation of LiDAR Point Clouds

RA-L 2021

Panoptic segmentation has recently unified semantic and instance segmentation, previously addressed separately, thus taking a step further towards creating more comprehensive and efficient perception systems. In this letter, we present Panoster, a novel proposal-free panoptic segmentation method for

Cited by 76SourceScholar
2021

SO-Pose: Exploiting Self-Occlusion for Direct 6D Pose Estimation

ICCV 2021poster

Directly regressing all 6 degrees-of-freedom (6DoF) for the object pose (i.e. the 3D rotation and translation) in a cluttered environment from a single RGB image is a challenging problem. While end-to-end methods have recently demonstrated promising results at high efficiency, they are still inferio…

Cited by 160PDFcodeScholar
2021

SceneGraphFusion: Incremental 3D Scene Graph Prediction From RGB-D Sequences

CVPR 2021poster

Scene graphs are a compact and explicit representation successfully used in a variety of 2D scene understanding tasks. This work proposes a method to build up semantic scene graphs from a 3D environment incrementally given a sequence of RGB-D frames. To this end, we aggregate PointNet features from…

Cited by 186PDFScholar
2021

Semantic Image Alignment for Vehicle Localization

IROS 2021poster

Accurate and reliable localization is a fundamental requirement for autonomous vehicles to use map information in higher-level tasks such as navigation or planning. In this paper, we present a novel approach to vehicle localization in dense semantic maps, including vectorized high-definition maps or…

Cited by 6SourceScholar
2021

Spotlight-Based 3D Instrument Guidance for Autonomous Task in Robot-Assisted Retinal Surgery

RA-L 2021

Retinal surgery is known to be a complicated and challenging task for an ophthalmologist even for retina specialists. Image guided robot-assisted intervention is among the novel and promising solutions that may enhance human capabilities during microsurgery. In this paper, a novel method is proposed

Cited by 21SourceScholar
2021

Unconditional Scene Graph Generation

ICCV 2021poster

Despite recent advancements in single-domain or single-object image generation, it is still challenging to generate complex scenes containing diverse, multiple objects and their interactions. Scene graphs, composed of nodes as objects and directed-edges as relationships among objects, offer an alter…

Cited by 35PDFScholar
2021

Unsupervised Traffic Scene Generation with Synthetic 3D Scene Graphs

IROS 2021poster

Image synthesis driven by computer graphics achieved recently a remarkable realism, yet synthetic image data generated this way reveals a significant domain gap with respect to real-world data. This is especially true in autonomous driving scenarios, which represent a critical aspect for over-coming…

Cited by 12SourceScholar
2020

6D Camera Relocalization in Ambiguous Scenes via Continuous Multimodal Inference

ECCV 2020poster

We present a multimodal camera relocalization framework that captures ambiguities and uncertainties with continuous mixture models defined on the manifold of camera poses. In highly ambiguous environments, which can easily arise due to symmetries and repetitive structures in the scene, computing one…

2020

Automatic Normal Positioning of Robotic Ultrasound Probe Based Only on Confidence Map Optimization and Force Measurement

RA-L 2020

Acquiring good image quality is one of the main challenges for fully-automatic robot-assisted ultrasound systems (RUSS). The presented method aims at overcoming this challenge for orthopaedic applications by optimizing the orientation of the robotic ultrasound (US) probe, i.e. aligning the central a

Cited by 83SourceScholar
2020

Fairness by Learning Orthogonal Disentangled Representations

ECCV 2020poster

Learning discriminative powerful representations is a crucial step for machine learning systems. Introducing invariance against arbitrary nuisance or sensitive attributes while performing well on specific tasks is an important problem in representation learning. This is mostly approached by purging…

Cited by 114SourcePDFScholar
2020

Force-Ultrasound Fusion: Bringing Spine Robotic-US to the Next "Level"

RA-L 2020

Spine injections are commonly performed in several clinical procedures. The localization of the target vertebral level (i.e. the position of a vertebra in a spine) is typically done by back palpation or under X-ray guidance, yielding either higher chances of procedure failure or exposure to ionizing

Cited by 40SourceScholar
2020

Learning 3D Semantic Scene Graphs From 3D Indoor Reconstructions

CVPR 2020poster

Scene understanding has been of high interest in computer vision. It encompasses not only identifying objects in a scene, but also their relationships within the given context. With this goal, a recent line of works tackles 3D semantic segmentation and scene layout prediction. In our work we focus o…

Cited by 262PDFScholar
2020

Reflective-AR Display: An Interaction Methodology for Virtual-to-Real Alignment in Medical Robotics

RA-L 2020

Robot-assisted minimally invasive surgery has shown to improve patient outcomes, as well as reduce complications and recovery time for several clinical applications. While increasingly configurable robotic arms can maximize reach and avoid collisions in cluttered environments, positioning them appro

Cited by 8SourceScholar
2020

Self6D: Self-Supervised Monocular 6D Object Pose Estimation

ECCV 2020poster

6D object pose estimation is a fundamental problem in computer vision. Convolutional Neural Networks (CNNs) have recently proven to be capable of predicting reliable 6D pose estimates even from monocular images. Nonetheless, CNNs are identified as being extremely data-driven, and acquiring adequate…

2020

Semantic Image Manipulation Using Scene Graphs

CVPR 2020poster

Image manipulation can be considered a special case of image generation where the image to be produced is a modification of an existing image. Image generation and manipulation have been, for the most part, tasks that operate on raw pixels. However, the remarkable progress in learning rich image and…

Cited by 146PDFcodeScholar
2020

Signal Clustering With Class-Independent Segmentation

ICASSP 2020accepted

Radar signals have been dramatically increasing in complexity, limiting the source separation ability of traditional approaches. In this paper we propose a Deep Learning-based clustering method, which encodes concurrent signals into images, and, for the first time, tackles clustering with image segm…

Cited by 0SourceScholar
2020

SoftPoolNet: Shape Descriptor for Point Cloud Completion and Classification

ECCV 2020poster

Point clouds are often the default choice for many applications as they exhibit more flexibility and efficiency than volumetric data. Nevertheless, their unorganized nature - points are stored in an unordered way - makes them less suited to be processed by deep learning pipelines. In this paper, we…

Cited by 92SourcePDFScholar
2020

Structure-SLAM: Low-Drift Monocular SLAM in Indoor Environments

RA-L 2020

In this letter a low-drift monocular SLAM method is proposed targeting indoor scenarios, where monocular SLAM often fails due to the lack of textured surfaces. Our approach decouples rotation and translation estimation of the tracking process to reduce the long-term drift in indoor environments. In

Cited by 129SourceScholar
2020

Towards Unsupervised Learning for Instrument Segmentation in Robotic Surgery with Cycle-Consistent Adversarial Networks

IROS 2020poster

Surgical tool segmentation in endoscopic images is an important problem: it is a crucial step towards full instrument pose estimation and it is used for integration of pre- and intra-operative images into the endoscopic view. While many recent approaches based on convolutional neural networks have s…

Cited by 26SourceScholar
2020

Ultrasound-Guided Robotic Navigation with Deep Reinforcement Learning

IROS 2020poster

In this paper we introduce the first reinforcement learning (RL) based robotic navigation method which utilizes ultrasound (US) images as an input. Our approach combines state-of-the-art RL techniques, specifically deep Q-networks (DQN) with memory buffers and a binary classifier for deciding when t…

Cited by 57SourceScholar
2019

Attention-based Lane Change Prediction

ICRA 2019poster

Lane change prediction of surrounding vehicles is a key building block of path planning. The focus has been on increasing the accuracy of prediction by posing it purely as a function estimation problem at the cost of model understandability. However, the efficacy of any lane change prediction model…

Cited by 60SourceScholar
2019

Crowd-sourced Semantic Edge Mapping for Autonomous Vehicles

IROS 2019poster

Highly accurate maps of the road infrastructure are a crucial cornerstone for self-driving cars to enable navigation in complex traffic scenarios. Traditional methods for creating detailed maps of road environments involve expensive survey vehicles that cannot keep up with the frequent changes in th…

Cited by 34SourceScholar
2019

Explaining the Ambiguity of Object Detection and 6D Pose From Visual Data

ICCV 2019poster

3D object detection and pose estimation from a single image are two inherently ambiguous problems. Oftentimes, objects appear similar from different viewpoints due to shape symmetries, occlusion and repetitive textures. This ambiguity in both detection and pose estimation means that an object instan…

Cited by 138PDFScholar
2019

ForkNet: Multi-Branch Volumetric Semantic Completion From a Single Depth Image

ICCV 2019poster

We propose a novel model for 3D semantic completion from a single depth image, based on a single encoder and three separate generators used to reconstruct different geometric and semantic representations of the original and completed scene, all sharing the same latent space. To transfer information…

Cited by 75PDFScholar
2019

Needle Localization for Robot-assisted Subretinal Injection based on Deep Learning

ICRA 2019poster

Subretinal injection is known to be a complicated task for ophthalmologists to perform, the main sources of difficulties are the fine anatomy of the retina, insufficient visual feedback, and high surgical precision. Image guided robot-assisted surgery is one of the promising solutions that bring sig…

Cited by 25SourceScholar
2019

RIO: 3D Object Instance Re-Localization in Changing Indoor Environments

ICCV 2019oral

In this work, we introduce the task of 3D object instance re-localization (RIO): given one or multiple objects in an RGB-D scan, we want to estimate their corresponding 6DoF poses in another 3D scan of the same environment taken at a later point in time. We consider RIO a particularly important task…

Cited by 171PDFcodeScholar
2019

Robotic Ultrasound for Catheter Navigation in Endovascular Procedures

IROS 2019poster

Endovascular procedures require real time visual feedback on the location of inserted catheters. This is currently achieved using X-ray fluoroscopy, which causes exposure to radiation. This study describes an alternative method using a robotic ultrasound system for catheter tracking and navigation i…

Cited by 34SourceScholar
2019

Sampling-Free Epistemic Uncertainty Estimation Using Approximated Variance Propagation

ICCV 2019oral

We present a sampling-free approach for computing the epistemic uncertainty of a neural network. Epistemic uncertainty is an important quantity for the deployment of deep neural networks in safety-critical applications, since it represents how much one can trust predictions on new data. Recently pro…

Cited by 183PDFcodeScholar
2019

Towards Unsupervised Image Captioning With Shared Multimodal Embeddings

ICCV 2019poster

Understanding images without explicit supervision has become an important problem in computer vision. In this paper, we address image captioning by generating language descriptions of scenes without learning from annotated pairs of images and their captions. The core component of our approach is a s…

Cited by 144PDFcodeScholar
2019

Variational Object-Aware 3-D Hand Pose From a Single RGB Image

RA-L 2019

We propose an approach to estimate the 3D pose of a human hand while grasping objects from a single RGB image. Our approach is based on a probabilistic model implemented with deep architectures, which is used for regressing, respectively, the 2D hand joints heat maps and the 3D hand joints coordinat

Cited by 12SourceScholar
2018

A Minimalist Approach to Type-Agnostic Detection of Quadrics in Point Clouds

CVPR 2018poster

This paper proposes a segmentation-free, automatic and efficient procedure to detect general geometric quadric forms in point clouds, where clutter and occlusions are inevitable. Our everyday world is dominated by man-made objects which are designed using 3D primitives (such as planes, cones, sphere…

Cited by 15SourcePDFScholar
2018

An Observer-Based Fusion Method Using Multicore Optical Shape Sensors and Ultrasound Images for Magnetically-Actuated Catheters

ICRA 2018poster

Minimally invasive surgery involves using flexible medical instruments such as endoscopes and catheters. Magnetically actuated catheters can provide improved steering precision over conventional catheters. However, besides the actuation method, an accurate tip position is required for precise contro…

Cited by 44SourceScholar
2018

Analyzing and Exploiting NARX Recurrent Neural Networks for Long-Term Dependencies

ICLR 2018workshop

Recurrent neural networks (RNNs) have achieved state-of-the-art performance on many diverse tasks, from machine translation to surgical activity recognition, yet training RNNs to capture long-term dependencies remains difficult. To date, the vast majority of successful RNN architectures alleviate th…

Cited by 36SourceScholar
2018

Distortion-Aware Convolutional Filters for Dense Prediction in Panoramic Images

ECCV 2018poster

There is a high demand of 3D data for 360° panoramic images and videos, pushed by the growing availability on the market of specialized hardware for both capturing (e.g., omnidirectional cameras) as well as visualizing in 3D (e.g., head mounted displays) panoramic images and videos. At the same time…

Cited by 221SourcePDFScholar
2018

Fully-Convolutional Point Networks for Large-Scale Point Clouds

ECCV 2018poster

This work proposes a general-purpose, fully-convolutional network architecture for efficiently processing large-scale 3D data. One striking characteristic of our approach is its ability to process unorganized 3D representations such as point clouds as input, then transforming them internally to orde…

2018

Guide Me: Interacting With Deep Networks

CVPR 2018poster

Interaction and collaboration between humans and intelligent machines has become increasingly important as machine learning methods move into real-world applications that involve end users. While much prior work lies at the intersection of natural language and vision, such as image captioning or ima…

Cited by 39SourcePDFScholar
2018

Human Motion Analysis with Deep Metric Learning

ECCV 2018poster

Effectively measuring the similarity between two human motions is necessary for several computer vision tasks such as gait analysis, person identification and action retrieval. Nevertheless, we believe that traditional approaches such as L2 distance or Dynamic Time Warping based on hand-crafted loca…

Cited by 68SourcePDFScholar
2018

Real-Time Fully Incremental Scene Understanding on Mobile Platforms

RA-L 2018

We propose an online RGB-D based scene understanding method for indoor scenes running in real time on mobile devices. First, we incrementally reconstruct the scene via simultaneous localization and mapping and compute a three-dimensional (3-D) geometric segmentation by fusing segments obtained from

Cited by 30SourceScholar
2018

Situation Assessment for Planning Lane Changes: Combining Recurrent Models and Prediction

ICRA 2018poster

One of the greatest challenges towards fully autonomous cars is the understanding of complex and dynamic scenes. Such understanding is needed for planning of maneuvers, especially those that are particularly frequent such as lane changes. While in recent years advanced driver-assistance systems have…

Cited by 41SourceScholar
2018

Towards Robotic Eye Surgery: Marker-Free, Online Hand-Eye Calibration Using Optical Coherence Tomography Images

RA-L 2018

Ophthalmic microsurgery is known to be a challenging operation, which requires very precise and dexterous manipulation. Image guided robot-assisted surgery is a promising solution that brings significant improvements in outcomes and reduces the physical limitations of human surgeons. However, this t

Cited by 38SourceScholar
2018

When Regression Meets Manifold Learning for Object Recognition and Pose Estimation

ICRA 2018poster

In this work, we propose a method for object recognition and pose estimation from depth images using convolutional neural networks. Previous methods addressing this problem rely on manifold learning to learn low dimensional viewpoint descriptors and employ them in a nearest neighbor search on an est…

Cited by 35SourceScholar
2017

CNN-SLAM: Real-Time Dense Monocular SLAM With Learned Depth Prediction

CVPR 2017spotlight

Given the recent advances in depth prediction from Convolutional Neural Networks (CNNs), this paper investigates how predicted depth maps from a deep neural network can be deployed for the goal of accurate and dense monocular reconstruction. We propose a method where CNN-predicted dense depth maps a…

Cited by 1024PDFcodeScholar
2017

Learning in an Uncertain World: Representing Ambiguity Through Multiple Hypotheses

ICCV 2017poster

Many prediction tasks contain uncertainty. In some cases, uncertainty is inherent in the task itself. In future prediction, for example, many distinct outcomes are equally valid. In other cases, uncertainty arises from the way data is labeled. For example, in object detection, many objects of intere…

Cited by 235PDFScholar
2017

Long Short-Term Memory Kalman Filters: Recurrent Neural Estimators for Pose Regularization

ICCV 2017poster

One-shot pose estimation for tasks such as body joint localization, camera pose estimation, and object tracking are generally noisy, and temporal filters have been extensively used for regularization. One of the most widely-used methods is the Kalman filter, which is both extremely simple and genera…

Cited by 237PDFScholar
2017

Real-Time 3D Model Tracking in Color and Depth on a Single CPU Core

CVPR 2017poster

We present a novel method to track 3D models in color and depth data. To this end, we introduce approximations that accelerate the state-of-the-art in region-based tracking by an order of magnitude while retaining similar accuracy. Furthermore, we show how the method can be made more robust in the p…

Cited by 51PDFScholar
2017

SSD-6D: Making RGB-Based 3D Detection and 6D Pose Estimation Great Again

ICCV 2017oral

We present a novel method for detecting 3D model instances and estimating their 6D poses from RGB data in a single shot. To this end, we extend the popular SSD paradigm to cover the full 6D pose space and train on synthetic model data only. Our approach competes or surpasses current state-of-the-art…

Cited by 1284PDFScholar
2016

Automatic force-compliant robotic ultrasound screening of abdominal aortic aneurysms

IROS 2016poster

Ultrasound (US) imaging is commonly employed for the diagnosis and staging of abdominal aortic aneurysms (AAA), mainly due to its non-invasiveness and high availability. High inter-operator variability and a lack of repeatability of current US image acquisition impair the implementation of extensive…

Cited by 125SourceScholar
2016

Confidence-driven control of an ultrasound probe: Target-specific acoustic window optimization

ICRA 2016

We propose a control framework to optimize the quality of robotic ultrasound imaging while tracking an anatomical target. We use a multitask approach to control the in-plane motion of a convex probe mounted on the end-effector of a robotic arm, based not only on the position of the target in the ima

Cited by 27SourceScholar
2016

Incremental scene understanding on dense SLAM

IROS 2016poster

We present an architecture for online, incremental scene modeling which combines a SLAM-based scene understanding framework with semantic segmentation and object pose estimation. The core of this approach comprises a probabilistic inference scheme that predicts semantic labels for object hypotheses…

Cited by 36SourceScholar
2016

Robotic ultrasound trajectory planning for volume of interest coverage

ICRA 2016

Medical robotic ultrasound offers potential to assist interventions, ease long-term monitoring and reduce operator dependency. Various techniques for remote control of ultrasound probes through telemanipulation systems have been presented in the past, however not exploiting the potential of fully au

Cited by 48SourceScholar
2016

Sensor substitution for video-based action recognition

IROS 2016poster

There are many applications where domain-specific sensing, such as accelerometers, kinematics, or force sensing, provide unique and important information for control or for analysis of motion. However, it is not always the case that these sensors can be deployed or accessed beyond laboratory environ…

Cited by 34SourceScholar
2016

Toward real-time 3D ultrasound registration-based visual servoing for interventional navigation

ICRA 2016

While intraoperative imaging is commonly used to guide surgical interventions, automatic robotic support for image-guided navigation has not yet been established in clinical routine. In this paper, we propose a novel visual servoing framework that combines, for the first time, full image-based 3D ul

Cited by 34SourceScholar
2016

Volumetric 3D Tracking by Detection

CVPR 2016spotlight

In this paper, we propose a new framework for 3D tracking by detection based on fully volumetric representations. On one hand, 3D tracking by detection has shown robust use in the context of interaction (Kinect) and surface tracking. On the other hand, volumetric representations have recently been p…

Cited by 40PDFScholar
2016

When 2.5D is not enough: Simultaneous reconstruction, segmentation and recognition on dense SLAM

ICRA 2016

While the main trend of 3D object recognition has been to infer object detection from single views of the scene - i.e., 2.5D data - this work explores the direction on performing object recognition on 3D data that is reconstructed from multiple viewpoints, under the conjecture that such data can imp

Cited by 95SourceScholar
2015

3D ultrasound-guided robotic steering of a flexible needle via visual servoing

ICRA 2015poster

We present a method for the three-dimensional (3D) steering of a flexible needle under 3D ultrasound guidance. The proposed solution is based on a duty-cycling visual servoing strategy we designed in a previous work, and on a new needle tracking algorithm for 3D ultrasound. The flexible needle model…

Cited by 63SourceScholar
2015

A Versatile Learning-Based 3D Temporal Tracker: Scalable, Robust, Online

ICCV 2015poster

This paper proposes a temporal tracking algorithm based on Random Forest that uses depth images to estimate and track the 3D pose of a rigid object in real-time. Compared to the state of the art aimed at the same goal, our algorithm holds important attributes such as high robustness against holes an…

Cited by 87PDFScholar
2015

Total Variation Regularization of Shape Signals

CVPR 2015poster

This paper introduces the concept of shape signals, i.e., series of shapes which have a natural temporal or spatial ordering, as well as a variational formulation for the regularization of these signals. The proposed formulation can be seen as the shape-valued generalization of the Rudin-Osher-Fatem…

Cited by 13SourcePDFScholar
2015

Toward User-Specific Tracking by Detection of Human Shapes in Multi-Cameras

CVPR 2015poster

Human shape tracking consists in fitting a template model to temporal sequences of visual observations. It usually comprises an association step, that finds correspondences between the model and the input data, and a deformation step, that fits the model to the observations given correspondences. Mo…

Cited by 19SourcePDFScholar
2015

Weakly-Supervised Structured Output Learning With Flexible and Latent Graphs Using High-Order Loss Functions

ICCV 2015poster

We introduce two new structured output models that use a latent graph, which is flexible in terms of the number of nodes and structure, where the training process minimises a high-order loss function using a weakly annotated training set. These models are developed in the context of microscopy imagi…

Cited by 13PDFScholar