← Search

Cornelia Fermüller

32 accepted papers

2026

First Frame Is the Place to Go for Video Content Customization

CVPR 2026

What role does the first frame play in video generation models? Traditionally, it's viewed as the spatial-temporal starting point of a video, merely a seed for subsequent animation. In this work, we reveal a fundamentally different perspective: video models implicitly treat the first frame as a conc

Cited by 0SourcecodeScholar
2026

NavMoE: Hybrid Model and Learning-Based Traversability Estimation for Local Navigation Via Mixture of Experts

ICRA 2026poster

This paper explores traversability estimation for robot navigation. A key bottleneck in traversability estimation lies in efficiently achieving reliable and robust predictions while accurately encoding both geometric and semantic information across diverse environments. We introduce Navigation via M…

2026

PolarDepth: Polarization-Guided Monocular Depth for Visual Odometry

RA-L 2026

Glass surfaces remain challenging for indoor robot perception. Depth sensors and RGB-only monocular depth estimation often fail because of reflections, refractions, and low-texture regions. To this end, we present PolarDepth, a polarization-enhanced monocular depth framework for glass-dominant envir

Cited by 0SourceScholar
2025

Air-FAR: Fast and Adaptable Routing for Aerial Navigation in Large-Scale Complex Unknown Environments

ICRA 2025

This paper presents a novel approach for realtime 3D navigation in large-scale complex environments by introducing a hierarchical 3D visibility graph (V-graph) and an efficient path search method. The proposed algorithm addresses the computational challenges of V-graph construction and shortest path

Cited by 0SourceScholar
2025

FeelAnyForce: Estimating Contact Force Feedback from Tactile Sensation for Vision-Based Tactile Sensors

ICRA 2025

In this paper, we tackle the problem of estimating 3D contact forces using vision-based tactile sensors. In particular, our goal is to estimate contact forces over a large range (up to 15 N) on any objects while generalizing across different vision-based tactile sensors. Thus, we collected a dataset

Cited by 14SourceScholar
2025

Learning Normal Flow Directly From Events

ICCV 2025poster

Event-based motion field estimation is an important task. However, current optical flow methods face challenges: learning-based approaches, often frame-based and relying on CNNs, lack cross-domain transferability, while model-based methods, though more robust, are less accurate. To address the limit…

2025

Search-Based Path Planning in Interactive Environments Among Movable Obstacles

ICRA 2025

This paper investigates Path planning Among Movable Obstacles (PAMO), which seeks a minimum cost collision-free path among static obstacles from start to goal while allowing the robot to push away movable obstacles (i.e., objects) along its path when needed. To develop planners that are complete and

Cited by 2SourceScholar
2025

ViewActive: Active viewpoint optimization from a single image

IROS 2025

When observing objects, humans benefit from their spatial visualization and mental rotation ability to envision potential optimal viewpoints based on the current observation. This capability is crucial for enabling robots to achieve efficient and robust scene perception during operation, as optimal

Cited by 4SourcecodeScholar
2024

AcTExplore: Active Tactile Exploration on Unknown Objects

ICRA 2024poster

Tactile exploration plays a crucial role in understanding object structures for fundamental robotics tasks such as grasping and manipulation. However, efficiently exploring such objects using tactile sensors is challenging, primarily due to the large-scale unknown environments and limited sensing co…

Cited by 5SourceScholar
2024

Active Human Pose Estimation via an Autonomous UAV Agent

IROS 2024poster

One of the core activities of an active observer involves moving to secure a "better" view of the scene, where the definition of "better" is task-dependent. This paper focuses on the task of human pose estimation from videos capturing a person’s activity. Self-occlusions within the scene can complic…

Cited by 2SourceScholar
2024

MARVIS: Motion & Geometry Aware Real and Virtual Image Segmentation

IROS 2024poster

Tasks such as autonomous navigation, 3D reconstruction, and object recognition near the water surfaces are crucial in marine robotics applications. However, challenges arise due to dynamic disturbances, e.g., light reflections and refraction from the random air-water interface, irregular liquid flow…

Cited by 3SourcecodeScholar
2023

TTCDist: Fast Distance Estimation From an Active Monocular Camera Using Time-to-Contact

ICRA 2023poster

Distance estimation from vision is fundamental for a myriad of robotic applications such as navigation, manipu-lation, and planning. Inspired by the mammal's visual system, which gazes at specific objects, we develop two novel constraints relating time-to-contact, acceleration, and distance that we…

Cited by 6SourceScholar
2023

Therbligs in Action: Video Understanding Through Motion Primitives

CVPR 2023poster

In this paper we introduce a rule-based, compositional, and hierarchical modeling of action using Therbligs as our atoms. Introducing these atoms provides us with a consistent, expressive, contact-centered representation of action. Over the atoms we introduce a differentiable method of rule-based re…

Cited by 17SourcePDFScholar
2023

WorldGen: A Large Scale Generative Simulator

ICRA 2023poster

In the era of deep learning, data is the critical determining factor in the performance of neural network models. Generating large datasets suffers from various challenges such as scalability, cost efficiency and photorealism. To avoid expensive and strenuous dataset collection and annotations, rese…

Cited by 7SourceScholar
2022

DiffPoseNet: Direct Differentiable Camera Pose Estimation

CVPR 2022poster

Current deep neural network approaches for camera pose estimation rely on scene structure for 3D motion estimation, but this decreases the robustness and thereby makes cross-dataset generalization difficult. In contrast, classical approaches to structure from motion estimate 3D motion utilizing opti…

Cited by 40PDFScholar
2021

0-MMS: Zero-Shot Multi-Motion Segmentation With A Monocular Event Camera

ICRA 2021poster

Segmentation of moving objects in dynamic scenes is a key process in scene understanding for navigation tasks. Classical cameras suffer from motion blur in such scenarios rendering them effete. On the contrary, event cameras, because of their high temporal resolution and lack of motion blur, are tai…

Cited by 38SourcecodeScholar
2021

EVPropNet: Detecting Drones By Finding Propellers For Mid-Air Landing And Following

RSS 2021poster

The rapid rise of accessibility of unmanned aerial vehicles or drones pose a threat to general security and confidentiality. Most of the commercially available or custom-built drones are multi-rotors and are comprised of multiple propellers. Since these propellers rotate at a high-speed; they are ge…

Cited by 17SourcePDFScholar
2021

MorphEyes: Variable Baseline Stereo For Quadrotor Navigation

ICRA 2021poster

Morphable design and depth-based visual control are two upcoming trends leading to advancements in the field of quadrotor autonomy. Stereo-cameras have struck the perfect balance of weight and accuracy of depth estimation but suffer from the problem of depth range being limited and dictated by the b…

Cited by 14SourcecodeScholar
2021

NudgeSeg: Zero-Shot Object Segmentation by Repeated Physical Interaction

IROS 2021poster

Recent advances in object segmentation have demonstrated that deep neural networks excel at object segmentation for specific classes in color and depth images. However, their performance is dictated by the number of classes and objects used for training, thereby hindering generalization to never see…

Cited by 5SourceScholar
2021

SpikeMS: Deep Spiking Neural Network for Motion Segmentation

IROS 2021poster

Spiking Neural Networks (SNN) are the so-called third generation of neural networks which attempt to more closely match the functioning of the biological brain. They inherently encode temporal data, allowing for training with less energy usage and can be extremely energy efficient when coded on neur…

Cited by 42SourceScholar
2020

EVDodgeNet: Deep Dynamic Obstacle Dodging with Event Cameras

ICRA 2020poster

Dynamic obstacle avoidance on quadrotors requires low latency. A class of sensors that are particularly suitable for such scenarios are event cameras. In this paper, we present a deep learning based solution for dodging multiple dynamic obstacles on a quadrotor with a single event camera and on-boar…

Cited by 96SourcecodeScholar
2020

Unsupervised Learning of Dense Optical Flow, Depth and Egomotion with Event-Based Sensors

IROS 2020poster

We present an unsupervised learning pipeline for dense depth, optical flow and egomotion estimation for autonomous driving applications, using the event-based output of the Dynamic Vision Sensor (DVS) as input. The backbone of our pipeline is a bioinspired encoder-decoder neural network architecture…

Cited by 73SourceScholar
2019

EV-IMO: Motion Segmentation Dataset and Learning Pipeline for Event Cameras

IROS 2019poster

We present the first event-based learning approach for motion segmentation in indoor scenes and the first event-based dataset - EV-IMO- which includes accurate pixel-wise motion masks, egomotion and ground truth depth. Our approach is based on an efficient implementation of the SfM learning pipeline…

Cited by 121SourceScholar
2018

Event-Based Moving Object Detection and Tracking

IROS 2018poster

Event-based vision sensors, such as the Dynamic Vision Sensor (DVS), are ideally suited for real-time motion analysis. The unique properties encompassed in the readings of such sensors provide high temporal resolution, superior sensitivity to light and low latency. These properties provide the groun…

Cited by 406SourceScholar
2018

GapFlyt: Active Vision Based Minimalist Structure-Less Gap Detection For Quadrotor Flight

RA-L 2018

Although quadrotors, and aerial robots in general, are inherently active agents, their perceptual capabilities in literature so far have been mostly passive in nature. Researchers and practitioners today use traditional computer vision algorithms with the aim of building a representation of general

Cited by 88SourcecodeScholar
2018

Seeing Behind the Scene: Using Symmetry to Reason About Objects in Cluttered Environments

IROS 2018poster

Symmetry is a common property shared by the majority of man-made objects. This paper presents a novel bottom-up approach for segmenting symmetric objects and recovering their symmetries from 3D pointclouds of natural scenes. Candidate rotational and reflectional symmetries are detected by fitting sy…

Cited by 13SourceScholar
2017

Fast task-specific target detection via graph based constraints representation and checking

ICRA 2017poster

We present a framework for fast target detection in real-world robotics applications. Considering that an intelligent agent attends to a task-specific object target during execution, our goal is to detect the object efficiently. We propose the concept of early recognition, which influences the candi…

Cited by 1SourceScholar
2017

What can i do around here? Deep functional scene understanding for cognitive robots

ICRA 2017poster

For robots that have the capability to interact with the physical environment through their end effectors, understanding the surrounding scenes is not merely a task of image classification or object recognition. To perform actual tasks, it is critical for the robot to have a functional understanding…

Cited by 59SourceScholar
2015

Affordance detection of tool parts from geometric features

ICRA 2015poster

As robots begin to collaborate with humans in everyday workspaces, they will need to understand the functions of tools and their parts. To cut an apple or hammer a nail, robots need to not just know the tool's name, but they must localize the tool's parts and identify their functions. Intuitively, t…

Cited by 386SourceScholar
2015

Learning the spatial semantics of manipulation actions through preposition grounding

ICRA 2015poster

In this paper, we introduce an abstract representation for manipulation actions that is based on the evolution of the spatial relations between involved objects. Object tracking in RGBD streams enables straightforward and intuitive ways to model spatial relations in 3D space. Reasoning in 3D overcom…

Cited by 64SourceScholar