← Search

Feras Dayoub

37 accepted papers

2025

3D-LLaVA: Towards Generalist 3D LMMs with Omni Superpoint Transformer

CVPR 2025poster

Current 3D Large Multimodal Models (3D LMMs) have shown tremendous potential in 3D-vision-based dialogue and reasoning. However, how to further enhance 3D LMMs to achieve fine-grained scene understanding and facilitate flexible human-agent interaction remains a challenging problem. In this work, we…

2025

Effective Tuning Strategies for Generalist Robot Manipulation Policies

ICRA 2025

Generalist robot manipulation policies (GMPs) have the potential to generalize across a wide range of tasks, devices, and environments. However, existing policies continue to struggle with out-of-distribution scenarios due to the inherent difficulty of collecting sufficient action data to cover exte

Cited by 9SourceScholar
2025

ObjectReact: Learning Object-Relative Control for Visual Navigation

CoRL 2025poster

Visual navigation using only a single camera and a topological map has recently become an appealing alternative to methods that require additional sensors and 3D maps. This is typically achieved through an "image-relative" approach to estimating control from a given pair of current observation and s…

Cited by 0SourceScholar
2025

QueryAdapter: Rapid Adaptation of Vision-Language Models in Response to Natural Language Queries

IROS 2025

A domain shift exists between the large-scale, internet data used to train a Vision-Language Model (VLM) and the raw image streams collected by a robot. Existing adaptation strategies require the definition of a closed-set of classes, which is impractical for a robot that must respond to diverse nat

Cited by 1SourceScholar
2025

Robust Scene Change Detection Using Visual Foundation Models and Cross-Attention Mechanisms

ICRA 2025

We present a novel method for scene change detection that leverages the robust feature extraction capabilities of a visual foundational model, DINOv2, and integrates full-image cross-attention to address key challenges such as varying lighting, seasonal variations, and viewpoint differences. In orde

Cited by 12SourcecodeScholar
2025

SmartWay: Enhanced Waypoint Prediction and Backtracking for Zero-Shot Vision-and-Language Navigation

IROS 2025

Vision-and-Language Navigation (VLN) in continuous environments requires agents to interpret natural language instructions while navigating unconstrained 3D spaces. Existing VLN-CE frameworks rely on a two-stage approach: a waypoint predictor to generate waypoints and a navigator to execute movement

Cited by 13SourceScholar
2025

TANGO: Traversability-Aware Navigation with Local Metric Control for Topological Goals

ICRA 2025

Visual navigation in robotics traditionally relies on globally-consistent 3D maps or learned controllers, which can be computationally expensive and difficult to generalize across diverse environments. In this work, we present a novel RGB-only, object-level topometric navigation pipeline that enable

Cited by 4SourcecodeScholar
2024

Physically Embodied Gaussian Splatting: A Visually Learnt and Physically Grounded 3D Representation for Robotics

CoRL 2024poster

For robots to robustly understand and interact with the physical world, it is highly beneficial to have a comprehensive representation -- modelling geometry, physics, and visual observations -- that informs perception, planning, and control algorithms. We propose a novel dual "Gaussian-Particle" re…

Cited by 9SourcecodeScholar
2024

RoboHop: Segment-based Topological Map Representation for Open-World Visual Navigation

ICRA 2024poster

Mapping is crucial for spatial reasoning, planning and robot navigation. Existing approaches range from metric, which require precise geometry-based optimization, to purely topological, where image-as-node based graphs lack explicit object-level reasoning and interconnectivity. In this paper, we pro…

Cited by 19SourcecodeScholar
2024

Segment beyond View: Handling Partially Missing Modality for Audio-Visual Semantic Segmentation

AAAI 2024technical

Augmented Reality (AR) devices, emerging as prominent mobile interaction platforms, face challenges in user safety, particularly concerning oncoming vehicles. While some solutions leverage onboard camera arrays, these cameras often have limited field-of-view (FoV) with front or downward perspectives…

Cited by 6SourcePDFScholar
2023

Predicting Class Distribution Shift for Reliable Domain Adaptive Object Detection

RA-L 2023

Unsupervised Domain Adaptive Object Detection (UDA-OD) uses unlabelled data to improve the reliability of robotic vision systems in open-world environments. Previous approaches to UDA-OD based on self-training have been effective in overcoming changes in the general appearance of images. However, sh

Cited by 11SourcecodeScholar
2023

SAFE: Sensitivity-Aware Features for Out-of-Distribution Object Detection

ICCV 2023poster

We address the problem of out-of-distribution (OOD) detection for the task of object detection. We show that residual convolutional layers with batch normalisation produce Sensitivity-Aware FEatures (SAFE) that are consistently powerful for distinguishing in-distribution from out-of-distribution det…

Cited by 37PDFcodeScholar
2022

FSNet: A Failure Detection Framework for Semantic Segmentation

RA-L 2022

Semantic segmentation is an important task that helps autonomous vehicles understand their surroundings and navigate safely. However, during deployment, even the most mature segmentation models are vulnerable to various external factors that can degrade the segmentation performance with potentially

Cited by 21SourceScholar
2022

Uncertainty for Identifying Open-Set Errors in Visual Object Detection

RA-L 2022

Deployed into an open world, object detectors are prone to open-set errors, false positive detections of object classes not present in the training dataset.We propose GMM-Det, a real-time method for extracting epistemic uncertainty from object detectors to identify and reject open-set errors. GMM-De

Cited by 54SourcecodeScholar
2021

Evaluating the Impact of Semantic Segmentation and Pose Estimation on Dense Semantic SLAM

IROS 2021poster

Recent Semantic SLAM methods combine classical geometry-based estimation with deep learning-based object detection or semantic segmentation. In this paper we evaluate the quality of semantic maps generated by state-of-the-art class-and instance-aware dense semantic SLAM algorithms whose codes are pu…

Cited by 11SourceScholar
2021

Online Monitoring of Object Detection Performance During Deployment

IROS 2021poster

During deployment, an object detector is expected to operate at a similar performance level reported on its testing dataset. However, when deployed onboard mobile robots that operate under varying and complex environmental conditions, the detector’s performance can fluctuate and occasionally degrade…

Cited by 13SourceScholar
2019

Did You Miss the Sign? A False Negative Alarm System for Traffic Sign Detectors

IROS 2019poster

Object detection is an integral part of an autonomous vehicle for its safety-critical and navigational purposes. Traffic sign as an object plays a vital role in guiding such systems. However, if the vehicle fails to locate any critical sign, it might make a catastrophic failure. In this paper, we ar…

Cited by 39SourceScholar
2019

Evaluating Merging Strategies for Sampling-based Uncertainty Techniques in Object Detection

ICRA 2019poster

There has been a recent emergence of sampling-based techniques for estimating epistemic uncertainty in deep neural networks. While these methods can be applied to classification or semantic segmentation tasks by simply averaging samples, this is not the case for object detection, where detection sam…

Cited by 140SourceScholar
2019

Predictive and adaptive maps for long-term visual navigation in changing environments

IROS 2019poster

In this paper, we compare different map management techniques for long-term visual navigation in changing environments. In this scenario, the navigation system needs to continuously update and refine its feature map in order to adapt to the environment appearance change. To achieve reliable long-ter…

Cited by 33SourceScholar
2018

Assisted Control for Semi-Autonomous Power Infrastructure Inspection Using Aerial Vehicles

IROS 2018poster

This paper presents the design and implementation of an assisted control technology for a small multirotor platform for aerial inspection of fixed energy infrastructure. Sensor placement is supported by a theoretical analysis of expected sensor performance and constrained platform behaviour to speed…

Cited by 5SourceScholar
2018

Dropout Sampling for Robust Object Detection in Open-Set Conditions

ICRA 2018poster

Dropout Variational Inference, or Dropout Sampling, has been recently proposed as an approximation technique for Bayesian Deep Learning and evaluated for image classification and regression tasks. This paper investigates the utility of Dropout Sampling for object detection for the first time. We dem…

Cited by 305SourceScholar
2018

Introduction to the Special Issue on Precision Agricultural Robotics and Autonomous Farming Technologies

RA-L 2018

Growth in world population, increasing urbanization and changing consumption habits mean demand for food production is predicted to increase dramatically over the coming decades. This increased demand for food production must be achieved despite challenges such as climate change, a limited supply of

Cited by 5SourceScholar
2017

A transplantable system for weed classification by agricultural robotics

IROS 2017poster

This work presents a rapidly deployable system for automated precision weeding with minimal human labeling time. This overcomes a limiting factor in robotic precision weeding related to the use of vision-based classification systems trained for species that may not be relevant to specific farms. We…

Cited by 6SourceScholar
2017

Peduncle Detection of Sweet Pepper for Autonomous Crop Harvesting - Combined Color and 3-D Information

RA-L 2017

This letter presents a three-dimensional (3-D) visual detection method for the challenging task of detecting peduncles of sweet peppers (Capsicum annuum) in the field. Cutting the peduncle cleanly is one of the most difficult stages of the harvesting process, where the peduncle is the part of the cr

Cited by 124SourceScholar
2016

Find my office: Navigating real space from semantic descriptions

ICRA 2016

This paper shows that by using only symbolic language phrases, a mobile robot can purposefully navigate to specified rooms in previously unexplored environments. The robot intelligently organises a symbolic language description of the unseen environment and “imagines” a representative map, called th

Cited by 19SourceScholar
2016

Place categorization and semantic mapping on a mobile robot

ICRA 2016

In this paper we focus on the challenging problem of place categorization and semantic mapping on a robot without environment-specific training. Motivated by their ongoing success in various visual recognition tasks, we build our system upon a state-of-the-art convolutional network. We overcome its

Cited by 143SourceScholar
2016

Visual detection of occluded crop: For automated harvesting

ICRA 2016

This paper presents a novel crop detection system applied to the challenging task of field sweet pepper (capsicum) detection. The field-grown sweet pepper crop presents several challenges for robotic systems such as the high degree of occlusion and the fact that the crop can have a similar colour to

Cited by 77SourceScholar
2015

On the performance of ConvNet features for place recognition

IROS 2015poster

After the incredible success of deep learning in the computer vision domain, there has been much interest in applying Convolutional Network (ConvNet) features in robotic fields such as visual navigation and SLAM. Unfortunately, there are fundamental differences and challenges involved. Computer visi…

Cited by 683SourceScholar
2015

Place Recognition with ConvNet Landmarks: Viewpoint-Robust, Condition-Robust, Training-Free

RSS 2015poster

Place recognition has long been an incompletely solved problem in that all approaches involve significant com- promises. Current methods address many but never all of the critical challenges of place recognition _ viewpoint-invariance, condition-invariance and minimizing training requirements. Here…

Cited by 503SourcePDFScholar
2015

Robot navigation using human cues: A robot navigation system for symbolic goal-directed exploration

ICRA 2015poster

In this paper we present for the first time a complete symbolic navigation system that performs goal-directed exploration to unfamiliar environments on a physical robot. We introduce a novel construct called the abstract map to link provided symbolic spatial information with observed symbolic inform…

Cited by 38SourceScholar