← Search

Alexander Hermans

14 accepted papers

2025

OoDIS: Anomaly Instance Segmentation and Detection Benchmark

ICRA 2025

Safe navigation of self-driving cars and robots requires a precise understanding of their environment. Training data for perception systems cannot cover the wide variety of objects that may appear during deployment. Thus, reliable identification of unknown objects, such as wild animals and untypical

Cited by 7SourceScholar
2025

Your ViT is Secretly an Image Segmentation Model

CVPR 2025highlight

Vision Transformers (ViTs) have shown remarkable performance and scalability across various computer vision tasks. To apply single-scale ViTs to image segmentation, existing methods adopt a convolutional adapter to generate multi-scale features, a pixel decoder to fuse these features, and a Transfor…

2023

DynaMITe: Dynamic Query Bootstrapping for Multi-object Interactive Segmentation Transformer

ICCV 2023poster

Most state-of-the-art instance segmentation methods rely on large amounts of pixel-precise ground-truth annotations for training, which are expensive to create. Interactive segmentation networks help generate such annotations based on an image and the corresponding user interactions such as clicks.…

Cited by 12PDFScholar
2023

Mask3D: Mask Transformer for 3D Semantic Instance Segmentation

ICRA 2023poster

Modern 3D semantic instance segmentation approaches predominantly rely on specialized voting mechanisms followed by carefully designed geometric clustering techniques. Building on the successes of recent Transformer-based methods for object detection and image segmentation, we propose the first Tran…

Cited by 272SourcecodeScholar
2023

TarViS: A Unified Approach for Target-Based Video Segmentation

CVPR 2023highlight

The general domain of video segmentation is currently fragmented into different tasks spanning multiple benchmarks. Despite rapid progress in the state-of-the-art, current methods are overwhelmingly task-specific and cannot conceptually generalize to other tasks. Inspired by recent approaches with m…

2022

HODOR: High-Level Object Descriptors for Object Re-Segmentation in Video Learned From Static Images

CVPR 2022oral

Existing state-of-the-art methods for Video Object Segmentation (VOS) learn low-level pixel-to-pixel correspondences between frames to propagate object masks across video. This requires a large amount of densely annotated video data, which is costly to annotate, and largely redundant since frames wi…

Cited by 30PDFcodeScholar
2021

Self-Supervised Person Detection in 2D Range Data using a Calibrated Camera

ICRA 2021poster

Deep learning is the essential building block of state-of-the-art person detectors in 2D range data. However, only a few annotated datasets are available for training and testing these deep networks, potentially limiting their performance when deployed in new environments or with different LiDAR mod…

Cited by 18SourcecodeScholar
2020

DR-SPAAM: A Spatial-Attention and Auto-regressive Model for Person Detection in 2D Range Data

IROS 2020poster

Detecting persons using a 2D LiDAR is a challenging task due to the low information content of 2D range data. To alleviate the problem caused by the sparsity of the LiDAR points, current state-of-the-art methods fuse multiple previous scans and perform detection using the combined scans. The downsid…

Cited by 45SourcecodeScholar
2018

Deep Person Detection in Two-Dimensional Range Data

RA-L 2018

Detecting humans is a key skill for mobile robots and intelligent vehicles in a large variety of applications. Although the problem is well studied for certain sensory modalities such as image data, few works exist that address this detection task using two-dimensional (2-D) range data. However, a w

Cited by 22SourceScholar
2018

MaskLab: Instance Segmentation by Refining Object Detection With Semantic and Direction Features

CVPR 2018poster

In this work, we tackle the problem of instance segmentation, the task of simultaneously solving object detection and semantic segmentation. Towards this goal, we present a model, called MaskLab, which produces three outputs: box detection, semantic segmentation, and direction prediction. Building o…

Cited by 497SourcePDFScholar
2017

Full-Resolution Residual Networks for Semantic Segmentation in Street Scenes

CVPR 2017oral

Semantic image segmentation is an essential component of modern autonomous driving systems, as an accurate understanding of the surrounding scene is crucial to navigation and action planning. Current state-of-the-art approaches in semantic image segmentation rely on pre-trained networks that were in…

Cited by 741PDFcodeScholar
2016

Multi-scale object candidates for generic object tracking in street scenes

ICRA 2016

Most vision based systems for object tracking in urban environments focus on a limited number of important object categories such as cars or pedestrians, for which powerful detectors are available. However, practical driving scenarios contain many additional objects of interest, for which suitable d

Cited by 44SourceScholar