← Search

Rainer Stiefelhagen

58 accepted papers

2026

Exploring Single Domain Generalization of LiDAR-Based Semantic Segmentation under Imperfect Labels

ICRA 2026poster

Accurate perception is critical for vehicle safety, with LiDAR as a key enabler in autonomous driving. To ensure robust performance across environments, sensor types, and weather conditions without costly re-annotation, domain generalization in LiDAR-based 3D semantic segmentation is essential. Howe…

2026

Go Beyond Earth: Understanding Human Actions and Scenes in Microgravity Environments

ICLR 2026poster

Despite substantial progress in video understanding, most existing datasets are limited to Earth’s gravitational conditions. However, microgravity alters human motion, interactions, and visual semantics, revealing a critical gap for real-world vision systems. This presents a challenge for domain-rob…

Cited by 0SourcecodeScholar
2026

HybriDLA: Hybrid Generation for Document Layout Analysis

AAAI 2026technical

Conventional document layout analysis (DLA) traditionally depends on empirical priors or a fixed set of learnable queries executed in a single forward pass. While sufficient for early-generation documents with a small, predetermined number of regions, this paradigm struggles with contemporary docume

Cited by 0SourcePDFScholar
2026

MICA: Multi-Agent Industrial Coordination Assistant

ICRA 2026poster

Industrial workflows demand adaptive and trustworthy assistance that can operate under limited computing, connectivity, and strict privacy constraints. In this work, we present MICA (Multi-Agent Industrial Coordination Assistant), a perception-grounded and speech-interactive system that delivers rea…

2026

More than the Sum: Panorama-Language Models for Adverse Omni-Scenes

CVPR 2026

Existing vision-language models (VLMs) are tailored for pinhole imagery, stitching multiple narrow field-of-view inputs to piece together a complete omni-scene understanding. Yet, such multi-view perception overlooks the holistic spatial and contextual relationships that a single panorama inherently

Cited by 0SourcecodeScholar
2026

RHO: Robust Holistic OSM-Based Metric Cross-View Geo-Localization

CVPR 2026

Metric Cross-View Geo-Localization (MCVGL) aims to estimate the 3-DoF camera pose (position and heading) by matching ground and satellite images. In this work, instead of pinhole and satellite images, we study robust MCVGL using holistic panoramas and OpenStreetMap (OSM). To this end, we establish a

Cited by 0SourcecodeScholar
2025

Every Component Counts: Rethinking the Measure of Success for Medical Semantic Segmentation in Multi-Instance Segmentation Tasks

AAAI 2025technical

We present Connected-Component (CC)-Metrics, a novel semantic segmentation evaluation protocol, targeted to align existing semantic segmentation metrics to a multi-instance detection scenario in which each connected component matters. We motivate this setup in the common medical scenario of semantic…

2025

Graph-based Document Structure Analysis

ICLR 2025poster

When reading a document, glancing at the spatial layout of a document is an initial step to understand it roughly. Traditional document layout analysis (DLA) methods, however, offer only a superficial parsing of documents, focusing on basic instance detection and often failing to capture the nuanced…

Cited by 0SourcePDFScholar
2025

HopaDIFF: Holistic-Partial Aware Fourier Conditioned Diffusion for Referring Human Action Segmentation in Multi-Person Scenarios

NeurIPS 2025spotlight

Action segmentation is a core challenge in high-level video understanding, aiming to partition untrimmed videos into segments and assign each a label from a predefined action set. Existing methods primarily address single-person activities with fixed action sequences, overlooking multi-person scenar…

Cited by 0SourcecodeScholar
2025

Is Visual in-Context Learning for Compositional Medical Tasks within Reach?

ICCV 2025poster

In this paper, we explore the potential of visual in-context learning to enable a single model to handle multiple tasks and adapt to new tasks during test time without re-training. Unlike previous approaches, our focus is on training in-context learners to adapt to sequences of tasks, rather than in…

2025

Scene-agnostic Pose Regression for Visual Localization

CVPR 2025poster

Absolute Pose Regression (APR) predicts 6D camera poses but lacks the adaptability to unknown environments without retraining, while Relative Pose Regression (RPR) generalizes better yet requires a large image retrieval database. Visual Odometry (VO) generalizes well in unseen environments but suffe…

Cited by 0SourcePDFScholar
2025

Situat3DChange: Situated 3D Change Understanding Dataset for Multimodal Large Language Model

NeurIPS 2025poster

Physical environments and circumstances are fundamentally dynamic, yet current 3D datasets and evaluation benchmarks tend to concentrate on either dynamic scenarios or dynamic situations in isolation, resulting in incomplete comprehension. To overcome these constraints, we introduce Situat3DChange,…

Cited by 0SourcecodeScholar
2025

VISO-Grasp: Vision-Language Informed Spatial Object-centric 6-DoF Active View Planning and Grasping in Clutter and Invisibility

IROS 2025

We propose VISO-Grasp, a novel vision-language-informed system designed to systematically address visibility constraints for grasping in severely occluded environments. By leveraging Foundation Models (FMs) for spatial reasoning and active view planning, our framework constructs and updates an insta

Cited by 13SourcecodeScholar
2025

mmWalk: Towards Multi-modal Multi-view Walking Assistance

NeurIPS 2025poster

Walking assistance in extreme or complex environments remains a significant challenge for people with blindness or low vision (BLV), largely due to the lack of a holistic scene understanding. Motivated by the real-world needs of the BLV community, we build mmWalk, a simulated multi-modal dataset tha…

Cited by 0SourceScholar
2024

Advancing Open-Set Domain Generalization Using Evidential Bi-Level Hardest Domain Scheduler

NeurIPS 2024poster

In Open-Set Domain Generalization (OSDG), the model is exposed to both new variations of data appearance (domains) and open-set conditions, where both known and novel categories are present at test time. The challenges of this task arise from the dual need to generalize across diverse domains and ac…

2024

Elevating Skeleton-Based Action Recognition with Efficient Multi-Modality Self-Supervision

ICASSP 2024accepted

Self-supervised representation learning for human action recognition has developed rapidly in recent years. Most of the existing works are based on skeleton data while using a multi-modality setup. These works overlooked the differences in performance among modalities, which led to the propagation o…

Cited by 0SourceScholar
2024

MateRobot: Material Recognition in Wearable Robotics for People with Visual Impairments

ICRA 2024poster

People with Visual Impairments (PVI) typically recognize objects through haptic perception. Knowing objects and materials before touching is desired by the target users but under-explored in the field of human-centered robotics. To fill this gap, in this work, a wearable vision-based robotic system,…

Cited by 13SourcecodeScholar
2024

Muscles in Time: Learning to Understand Human Motion In-Depth by Simulating Muscle Activations

NeurIPS 2024poster

Exploring the intricate dynamics between muscular and skeletal structures is pivotal for understanding human motion. This domain presents substantial challenges, primarily attributed to the intensive resources required for acquiring ground truth muscle activation data, resulting in a scarcity of dat…

Cited by 3SourcePDFScholar
2024

Navigating Open Set Scenarios for Skeleton-Based Action Recognition

AAAI 2024technical

In real-world scenarios, human actions often fall outside the distribution of training data, making it crucial for models to recognize known actions and reject unknown ones. However, using pure skeleton data in such open-set conditions poses challenges due to the lack of visual background cues and t…

2024

Occlusion-Aware Seamless Segmentation

ECCV 2024poster

"Panoramic images can broaden the Field of View (FoV), occlusion-aware prediction can deepen the understanding of the scene, and domain adaptation can transfer across viewing domains. In this work, we introduce a novel task, Occlusion-Aware Seamless Segmentation (OASS), which simultaneously tackles…

2024

Open Panoramic Segmentation

ECCV 2024poster

"Panoramic images, capturing a 360° field of view (FoV), encompass omnidirectional spatial information crucial for scene understanding. However, it is not only costly to obtain training-sufficient dense-annotated panoramas but also application-restricted when training models in a close-vocabulary se…

2024

Referring Atomic Video Action Recognition

ECCV 2024poster

"We introduce a new task called Referring Atomic Video Action Recognition (RAVAR), aimed at identifying atomic actions of a particular person based on a textual description and the video data of this person. This task differs from traditional action recognition and localization, where predictions ar…

2024

RoDLA: Benchmarking the Robustness of Document Layout Analysis Models

CVPR 2024poster

Before developing a Document Layout Analysis (DLA) model in real-world applications conducting comprehensive robustness testing is essential. However the robustness of DLA models remains underexplored in the literature. To address this we are the first to introduce a robustness benchmark for DLA mod…

Cited by 6SourcePDFScholar
2024

SciEx: Benchmarking Large Language Models on Scientific Exams with Human Expert Grading and Automatic Grading

EMNLP 2024main

With the rapid development of Large Language Models (LLMs), it is crucial to have benchmarks which can evaluate the ability of LLMs on different domains. One common use of LLMs is performing tasks on scientific topics, such as writing algorithms, querying databases or giving mathematical proofs. Ins…

2024

Skeleton-Based Human Action Recognition with Noisy Labels

IROS 2024poster

Understanding human actions from body poses is critical for assistive robots sharing space with humans in order to make informed and safe decisions about the next interaction. However, precise temporal localization and annotation of activity sequences is time-consuming and the resulting labels are o…

Cited by 5SourcecodeScholar
2024

Statewide Visual Geolocalization in the Wild

ECCV 2024poster

This work presents a method that is able to predict the geolocation of a street-view photo taken in the wild within a state-sized search region by matching against a database of aerial reference imagery. We partition the search region into geographical cells and train a model to map cells and corres…

2024

SynthAct: Towards Generalizable Human Action Recognition based on Synthetic Data

ICRA 2024poster

Synthetic data generation is a proven method for augmenting training sets without the need for extensive setups, yet its application in human activity recognition is underexplored. This is particularly crucial for human-robot collaboration in household settings, where data collection is often privac…

Cited by 3SourceScholar
2023

Decoupled Semantic Prototypes Enable Learning From Diverse Annotation Types for Semi-Weakly Segmentation in Expert-Driven Domains

CVPR 2023poster

A vast amount of images and pixel-wise annotations allowed our community to build scalable segmentation solutions for natural domains. However, the transfer to expert-driven domains like microscopy applications or medical healthcare remains difficult as domain experts are a critical factor due to th…

2023

Delivering Arbitrary-Modal Semantic Segmentation

CVPR 2023poster

Multimodal fusion can make semantic segmentation more robust. However, fusing an arbitrary number of modalities remains underexplored. To delve into this problem, we create the DeLiVER arbitrary-modal segmentation benchmark, covering Depth, LiDAR, multiple Views, Events, and RGB. Aside from this, we…

2023

Quantized Distillation: Optimizing Driver Activity Recognition Models for Resource-Constrained Environments

IROS 2023poster

Deep learning-based models are at the top of most driver observation benchmarks due to their remarkable accuracies but come with a high computational cost, while the resources are often limited in real-world driving scenarios. This paper presents a lightweight framework for resource- efficient drive…

Cited by 1SourcecodeScholar
2023

Uncertainty-Aware Vision-Based Metric Cross-View Geolocalization

CVPR 2023poster

This paper proposes a novel method for vision-based metric cross-view geolocalization (CVGL) that matches the camera images captured from a ground-based vehicle with an aerial image to determine the vehicle's geo-pose. Since aerial images are globally available at low cost, they represent a potentia…

Cited by 47SourcePDFScholar
2022

Bending Reality: Distortion-Aware Transformers for Adapting to Panoramic Semantic Segmentation

CVPR 2022poster

Panoramic images with their 360deg directional view encompass exhaustive information about the surrounding space, providing a rich foundation for scene understanding. To unfold this potential in the form of robust panoramic segmentation models, large quantities of expensive, pixel-wise annotations a…

Cited by 107PDFcodeScholar
2022

Continuous Self-Localization on Aerial Images Using Visual and Lidar Sensors

IROS 2022poster

This paper proposes a novel method for geo-tracking, i.e. continuous metric self-localization in outdoor environments by registering a vehicle's sensor information with aerial imagery of an unseen target region. Geo- tracking methods offer the potential to supplant noisy signals from global navigati…

Cited by 19SourceScholar
2022

Graph-Constrained Contrastive Regularization for Semi-Weakly Volumetric Segmentation

ECCV 2022poster

"Semantic volume segmentation suffers from the requirement of having voxel-wise annotated ground-truth data, which requires immense effort to obtain. In this work, we investigate how models can be trained from sparsely annotated volumes, i.e. volumes with only individual slices annotated. By formula…

2022

Hierarchical Nearest Neighbor Graph Embedding for Efficient Dimensionality Reduction

CVPR 2022poster

Dimensionality reduction is crucial both for visualization and preprocessing high dimensional data for machine learning. We introduce a novel method based on a hierarchy built on 1-nearest neighbor graphs in the original space which is used to preserve the grouping properties of the data distributio…

Cited by 23PDFcodeScholar
2022

Multimodal Generation of Novel Action Appearances for Synthetic-to-Real Recognition of Activities of Daily Living

IROS 2022poster

Domain shifts, such as appearance changes, are a key challenge in real-world applications of activity recognition models, which range from assistive robotics and smart homes to driver observation in intelligent vehicles. For example, while simulations are an excellent way of economical data collecti…

Cited by 3SourcecodeScholar
2022

Reference-Guided Pseudo-Label Generation for Medical Semantic Segmentation

AAAI 2022technical

Producing densely annotated data is a difficult and tedious task for medical imaging applications. To address this problem, we propose a novel approach to generate supervision for semi-supervised semantic segmentation. We argue that visually similar regions between labeled and unlabeled images lik…

Cited by 75SourcePDFScholar
2022

TransDARC: Transformer-based Driver Activity Recognition with Latent Space Feature Calibration

IROS 2022poster

Traditional video-based human activity recognition has experienced remarkable progress linked to the rise of deep learning, but this effect was slower as it comes to the downstream task of driver behavior understanding. Understanding the situation inside the vehicle cabin is essential for Advanced D…

Cited by 39SourcecodeScholar
2021

Capturing Omni-Range Context for Omnidirectional Segmentation

CVPR 2021poster

Convolutional Networks (ConvNets) excel at semantic segmentation and have become a vital component for perception in autonomous driving. Enabling an all-encompassing view of street-scenes, omnidirectional cameras present themselves as a perfect fit in such systems. Most segmentation models for parsi…

Cited by 91PDFcodeScholar
2021

Every Annotation Counts: Multi-Label Deep Supervision for Medical Image Segmentation

CVPR 2021poster

Pixel-wise segmentation is one of the most data and annotation hungry tasks in our field. Providing representative and accurate annotations is often mission-critical especially for challenging medical applications. In this paper, we propose a semi-weakly supervised segmentation algorithm to overcome…

Cited by 99PDFcodeScholar
2021

ISSAFE: Improving Semantic Segmentation in Accidents by Fusing Event-based Data

IROS 2021poster

Ensuring the safety of all traffic participants is a prerequisite for bringing intelligent vehicles closer to practical applications. The assistance system should not only achieve high accuracy under normal conditions, but obtain robust perception against extreme situations. However, traffic acciden…

Cited by 60SourcecodeScholar
2021

Let’s Play for Action: Recognizing Activities of Daily Living by Learning from Life Simulation Video Games

IROS 2021poster

Recognizing Activities of Daily Living (ADL) is a vital process for intelligent assistive robots, but collecting large annotated datasets requires time-consuming temporal labeling and raises privacy concerns, e.g., if the data is collected in a real household. In this work, we explore the concept of…

Cited by 49SourcecodeScholar
2021

Temporally-Weighted Hierarchical Clustering for Unsupervised Action Segmentation

CVPR 2021poster

Action segmentation refers to inferring boundaries of semantically consistent visual concepts in videos and is an important requirement for many video understanding tasks. For this and other video understanding tasks, supervised approaches have achieved encouraging performance but require a high vol…

Cited by 81PDFcodeScholar
2021

Vi2CLR: Video and Image for Visual Contrastive Learning of Representation

ICCV 2021poster

In this paper, we introduce a novel self-supervised visual representation learning method which understands both images and videos in a joint learning fashion. The proposed neural network architecture and objectives are designed to obtain two different Convolutional Neural Networks for solving visua…

Cited by 66PDFScholar
2020

Large Scale Holistic Video Understanding

ECCV 2020poster

Video recognition has been advanced in recent years by benchmarks with rich annotations. However, research is still mainly limited to human action or sports recognition - focusing on a highly specific video understanding task and thus leaving a significant gap towards describing the overall content…

2019

Drive&Act: A Multi-Modal Dataset for Fine-Grained Driver Behavior Recognition in Autonomous Vehicles

ICCV 2019poster

We introduce the novel domain-specific Drive&Act benchmark for fine-grained categorization of driver behavior. Our dataset features twelve hours and over 9.6 million frames of people engaged in distractive activities during both, manual and automated driving. We capture color, infrared, depth and 3D…

Cited by 243PDFScholar
2019

Efficient Parameter-Free Clustering Using First Neighbor Relations

CVPR 2019oral

We present a new clustering method in the form of a single clustering equation that is able to directly discover groupings in the data. The main proposition is that the first neighbor of each sample is all one needs to discover large chains and finding the groups in the data. In contrast to most exi…

Cited by 275PDFcodeScholar
2019

It's Not About the Journey; It's About the Destination: Following Soft Paths Under Question-Guidance for Visual Reasoning

CVPR 2019poster

Visual Reasoning remains a challenging task, as it has to deal with long-range and multi-step object relationships in the scene. We present a new model for Visual Reasoning, aimed at capturing the interplay among individual objects in the image represented as a scene graph. As not all graph componen…

Cited by 22PDFScholar
2018

3D Vehicle Trajectory Reconstruction in Monocular Video Data Using Environment Structure Constraints

ECCV 2018poster

We present a framework to reconstruct three-dimensional vehicle trajectories using monocular video data. We track two-dimensional vehicle shapes on pixel level exploiting instance-aware semantic segmentation techniques and optical flow cues. We apply Structure from Motion techniques to vehicle and b…

Cited by 15SourcePDFScholar
2018

A Pose-Sensitive Embedding for Person Re-Identification With Expanded Cross Neighborhood Re-Ranking

CVPR 2018poster

Person re-identification is a challenging retrieval task that requires matching a person’s acquired image across non-overlapping camera views. In this paper we propose an effective approach that incorporates both the fine and coarse pose information of the person to learn a discrim- inative embeddin…

Cited by 604SourcePDFScholar
2018

Classification-Driven Dynamic Image Enhancement

CVPR 2018poster

Convolutional neural networks rely on image texture and structure to serve as discriminative features to classify the image content. Image enhancement techniques can be used as preprocessing steps to help improve the overall image quality and in turn improve the overall effectiveness of a CNN. Exis…

Cited by 88SourcePDFScholar
2017

Automatic Discovery, Association Estimation and Learning of Semantic Attributes for a Thousand Categories

CVPR 2017poster

Attribute-based recognition models, due to their impressive performance and their ability to generalize well on novel categories, have been widely adopted for many computer vision applications. However, usually both the attribute vocabulary and the class-attribute associations have to be provided ma…

Cited by 38PDFScholar
2016

MovieQA: Understanding Stories in Movies Through Question-Answering

CVPR 2016spotlight

We introduce the MovieQA dataset which aims to evaluate automatic story comprehension from both video and text. The dataset consists of 14,944 questions about 408 movies with high semantic diversity. The questions range from simpler "Who" did "What" to "Whom", to "Why" and "How" certain events occur…

Cited by 875PDFScholar
2016

Recovering the Missing Link: Predicting Class-Attribute Associations for Unsupervised Zero-Shot Learning

CVPR 2016accepted

Collecting training images for all visual categories is not only expensive but also impractical. Zero-shot learning (ZSL), especially using attributes, offers a pragmatic solution to this problem. However, at test time most attribute-based methods require a full description of attribute associations…

Cited by 126SourcePDFScholar