← Search

Yiannis Aloimonos

57 accepted papers

2026

Adversarial Game-Theoretic Algorithm for Dexterous Grasp Synthesis

ICRA 2026poster

For many complex tasks, multi-finger robot hands are poised to revolutionize how we interact with the world, but reliably grasping objects remains a significant challenge. We focus on the problem of synthesizing grasps for multi-finger robot hands that, given an target object's geometry and pose, co…

2026

DREAM: Domain-Aware Reasoning for Efficient Autonomous Underwater Monitoring

ICRA 2026poster

The ocean is warming and acidifying, increasing the risk of mass mortality events for temperature-sensitive shellfish such as oysters. This motivates the development of long-term monitoring systems. However, human labor is costly and long-duration underwater work is highly hazardous, thus favoring r…

2026

First Frame Is the Place to Go for Video Content Customization

CVPR 2026

What role does the first frame play in video generation models? Traditionally, it's viewed as the spatial-temporal starting point of a video, merely a seed for subsequent animation. In this work, we reveal a fundamentally different perspective: video models implicitly treat the first frame as a conc

Cited by 0SourcecodeScholar
2026

From Inpainting to Layer Decomposition: Repurposing Generative Inpainting Models for Image Layer Decomposition

CVPR 2026

Images can be viewed as layered compositions, foreground objects over background, with potential occlusions. This layered representation enables independent editing of elements, offering greater flexibility for content creation. Despite the progress in large generative models, decomposing a single i

Cited by 0SourceScholar
2026

NavMoE: Hybrid Model and Learning-Based Traversability Estimation for Local Navigation Via Mixture of Experts

ICRA 2026poster

This paper explores traversability estimation for robot navigation. A key bottleneck in traversability estimation lies in efficiently achieving reliable and robust predictions while accurately encoding both geometric and semantic information across diverse environments. We introduce Navigation via M…

2026

PolarDepth: Polarization-Guided Monocular Depth for Visual Odometry

RA-L 2026

Glass surfaces remain challenging for indoor robot perception. Depth sensors and RGB-only monocular depth estimation often fail because of reflections, refractions, and low-texture regions. To this end, we present PolarDepth, a polarization-enhanced monocular depth framework for glass-dominant envir

Cited by 0SourceScholar
2025

Air-FAR: Fast and Adaptable Routing for Aerial Navigation in Large-Scale Complex Unknown Environments

ICRA 2025

This paper presents a novel approach for realtime 3D navigation in large-scale complex environments by introducing a hierarchical 3D visibility graph (V-graph) and an efficient path search method. The proposed algorithm addresses the computational challenges of V-graph construction and shortest path

Cited by 0SourceScholar
2025

Discovering Object Attributes by Prompting Large Language Models With Perception-Action Apis

ICRA 2025

There has been a lot of interest in grounding natural language to physical entities through visual context. While Vision Language Models (VLMs) can ground linguistic instructions to visual sensory information, they struggle with grounding non-visual attributes, like the weight of an object. Our key

Cited by 2SourceScholar
2025

FeelAnyForce: Estimating Contact Force Feedback from Tactile Sensation for Vision-Based Tactile Sensors

ICRA 2025

In this paper, we tackle the problem of estimating 3D contact forces using vision-based tactile sensors. In particular, our goal is to estimate contact forces over a large range (up to 15 N) on any objects while generalizing across different vision-based tactile sensors. Thus, we collected a dataset

Cited by 14SourceScholar
2025

Learning Normal Flow Directly From Events

ICCV 2025poster

Event-based motion field estimation is an important task. However, current optical flow methods face challenges: learning-based approaches, often frame-based and relying on CNNs, lack cross-domain transferability, while model-based methods, though more robust, are less accurate. To address the limit…

2025

ODYSSEE: Oyster Detection Yielded by Sensor Systems on Edge Electronics

ICRA 2025

Oysters are a vital keystone species in coastal ecosystems, providing significant economic, environmental, and cultural benefits. As the importance of oysters grows, so does the relevance of autonomous systems for their detection and monitoring. However, current monitoring strategies often rely on d

Cited by 5SourceScholar
2025

Repurposing Pre-trained Video Diffusion Models for Event-based Video Interpolation

CVPR 2025poster

Video Frame Interpolation aims to recover realistic missing frames between observed frames, generating a high-frame-rate video from a low-frame-rate video. However, without additional guidance, large motion between frames makes this problem ill-posed. Event-based Video Frame Interpolation (EVFI) add…

Cited by 4SourcePDFScholar
2025

ViewActive: Active viewpoint optimization from a single image

IROS 2025

When observing objects, humans benefit from their spatial visualization and mental rotation ability to envision potential optimal viewpoints based on the current observation. This capability is crucial for enabling robots to achieve efficient and robust scene perception during operation, as optimal

Cited by 4SourcecodeScholar
2024

A Linear Time and Space Local Point Cloud Geometry Encoder via Vectorized Kernel Mixture (VecKM)

ICML 2024poster

We propose VecKM, a local point cloud geometry encoder that is descriptive and efficient to compute. VecKM leverages a unique approach by vectorizing a kernel mixture to represent the local point cloud. Such representation's descriptiveness is supported by two theorems that validate its ability to r…

2024

AcTExplore: Active Tactile Exploration on Unknown Objects

ICRA 2024poster

Tactile exploration plays a crucial role in understanding object structures for fundamental robotics tasks such as grasping and manipulation. However, efficiently exploring such objects using tactile sensors is challenging, primarily due to the large-scale unknown environments and limited sensing co…

Cited by 5SourceScholar
2024

Active Human Pose Estimation via an Autonomous UAV Agent

IROS 2024poster

One of the core activities of an active observer involves moving to secure a "better" view of the scene, where the definition of "better" is task-dependent. This paper focuses on the task of human pose estimation from videos capturing a person’s activity. Self-occlusions within the scene can complic…

Cited by 2SourceScholar
2024

CodedEvents: Optimal Point-Spread-Function Engineering for 3D-Tracking with Event Cameras

CVPR 2024poster

Point-spread-function (PSF) engineering is a well-established computational imaging technique that uses phase masks and other optical elements to embed extra information (e.g. depth) into the images captured by conventional CMOS image sensors. To date however PSF-engineering has not been applied to…

Cited by 2SourcePDFScholar
2024

Cook2LTL: Translating Cooking Recipes to LTL Formulae using Large Language Models

ICRA 2024poster

Cooking recipes are challenging to translate to robot plans as they feature rich linguistic complexity, temporally-extended interconnected tasks, and an almost infinite space of possible actions. Our key insight is that combining a source of cooking domain knowledge with a formalism that captures th…

Cited by 23SourceScholar
2024

Decodable and Sample Invariant Continuous Object Encoder

ICLR 2024poster

We propose Hyper-Dimensional Function Encoding (HDFE). Given samples of a continuous object (e.g. a function), HDFE produces an explicit vector representation of the given object, invariant to the sample distribution and density. Sample distribution and density invariance enables HDFE to consistentl…

2024

Diving Deep into the Motion Representation of Video-Text Models

ACL 2024findings

Videos are more informative than images becausethey capture the dynamics of the scene.By representing motion in videos, we can capturedynamic activities. In this work, we introduceGPT-4 generated motion descriptions thatcapture fine-grained motion descriptions of activitiesand apply them to three ac…

2024

Embodiment: Self-Supervised Depth Estimation Based on Camera Models

IROS 2024poster

Depth estimationn is a critical topic for robotics and vision-related tasks. In monocular depth estimation, in comparison with supervised learning that requires expensive ground truth labeling, self-supervised methods possess great potential due to no labeling cost. However, self-supervised learning…

Cited by 1SourceScholar
2024

Event3DGS: Event-Based 3D Gaussian Splatting for High-Speed Robot Egomotion

CoRL 2024poster

By combining differentiable rendering with explicit point-based scene representations, 3D Gaussian Splatting (3DGS) has demonstrated breakthrough 3D reconstruction capabilities. However, to date 3DGS has had limited impact on robotics, where high-speed egomotion is pervasive: Egomotion introduc…

Cited by 11SourceScholar
2024

Interactive-FAR:Interactive, Fast and Adaptable Routing for Navigation Among Movable Obstacles in Complex Unknown Environments

IROS 2024poster

This paper introduces a real-time algorithm for navigating complex unknown environments cluttered with movable obstacles. Our algorithm achieves fast, adaptable routing by actively attempting to manipulate obstacles during path planning and adjusting the global plan from sensor feedback. The main co…

Cited by 3SourceScholar
2024

MARVIS: Motion & Geometry Aware Real and Virtual Image Segmentation

IROS 2024poster

Tasks such as autonomous navigation, 3D reconstruction, and object recognition near the water surfaces are crucial in marine robotics applications. However, challenges arise due to dynamic disturbances, e.g., light reflections and refraction from the random air-water interface, irregular liquid flow…

Cited by 3SourcecodeScholar
2024

Temporally Consistent Atmospheric Turbulence Mitigation with Neural Representations

NeurIPS 2024poster

Atmospheric turbulence, caused by random fluctuations in the atmosphere's refractive index, introduces complex spatio-temporal distortions in imagery captured at long range. Video Atmospheric Turbulence Mitigation (ATM) aims to restore videos affected by these distortions. However, existing video AT…

2024

UIVNAV: Underwater Information-driven Vision-based Navigation via Imitation Learning

ICRA 2024poster

Autonomous navigation in the underwater environment is challenging due to limited visibility, dynamic changes, and the lack of a cost-efficient, accurate localization system. We introduce UIVNAV, a novel end-to-end underwater navigation solution designed to navigate robots over Objects of Interest (…

Cited by 13SourceScholar
2023

Detecting Olives with Synthetic or Real Data? Olive the Above

IROS 2023poster

Modern robotics has enabled the advancement in yield estimation for precision agriculture. However, when applied to the olive industry, the high variation of olive colors and their similarity to the background leaf canopy presents a challenge. Labeling several thousands of very dense olive grove ima…

Cited by 2SourceScholar
2023

TTCDist: Fast Distance Estimation From an Active Monocular Camera Using Time-to-Contact

ICRA 2023poster

Distance estimation from vision is fundamental for a myriad of robotic applications such as navigation, manipu-lation, and planning. Inspired by the mammal's visual system, which gazes at specific objects, we develop two novel constraints relating time-to-contact, acceleration, and distance that we…

Cited by 6SourceScholar
2023

Therbligs in Action: Video Understanding Through Motion Primitives

CVPR 2023poster

In this paper we introduce a rule-based, compositional, and hierarchical modeling of action using Therbligs as our atoms. Introducing these atoms provides us with a consistent, expressive, contact-centered representation of action. Over the atoms we introduce a differentiable method of rule-based re…

Cited by 17SourcePDFScholar
2023

WorldGen: A Large Scale Generative Simulator

ICRA 2023poster

In the era of deep learning, data is the critical determining factor in the performance of neural network models. Generating large datasets suffers from various challenges such as scalability, cost efficiency and photorealism. To avoid expensive and strenuous dataset collection and annotations, rese…

Cited by 7SourceScholar
2022

DiffPoseNet: Direct Differentiable Camera Pose Estimation

CVPR 2022poster

Current deep neural network approaches for camera pose estimation rely on scene structure for 3D motion estimation, but this decreases the robustness and thereby makes cross-dataset generalization difficult. In contrast, classical approaches to structure from motion estimate 3D motion utilizing opti…

Cited by 40PDFScholar
2021

0-MMS: Zero-Shot Multi-Motion Segmentation With A Monocular Event Camera

ICRA 2021poster

Segmentation of moving objects in dynamic scenes is a key process in scene understanding for navigation tasks. Classical cameras suffer from motion blur in such scenarios rendering them effete. On the contrary, event cameras, because of their high temporal resolution and lack of motion blur, are tai…

Cited by 38SourcecodeScholar
2021

EVPropNet: Detecting Drones By Finding Propellers For Mid-Air Landing And Following

RSS 2021poster

The rapid rise of accessibility of unmanned aerial vehicles or drones pose a threat to general security and confidentiality. Most of the commercially available or custom-built drones are multi-rotors and are comprised of multiple propellers. Since these propellers rotate at a high-speed; they are ge…

Cited by 17SourcePDFScholar
2021

MorphEyes: Variable Baseline Stereo For Quadrotor Navigation

ICRA 2021poster

Morphable design and depth-based visual control are two upcoming trends leading to advancements in the field of quadrotor autonomy. Stereo-cameras have struck the perfect balance of weight and accuracy of depth estimation but suffer from the problem of depth range being limited and dictated by the b…

Cited by 14SourcecodeScholar
2021

NudgeSeg: Zero-Shot Object Segmentation by Repeated Physical Interaction

IROS 2021poster

Recent advances in object segmentation have demonstrated that deep neural networks excel at object segmentation for specific classes in color and depth images. However, their performance is dictated by the number of classes and objects used for training, thereby hindering generalization to never see…

Cited by 5SourceScholar
2021

SpikeMS: Deep Spiking Neural Network for Motion Segmentation

IROS 2021poster

Spiking Neural Networks (SNN) are the so-called third generation of neural networks which attempt to more closely match the functioning of the biological brain. They inherently encode temporal data, allowing for training with less energy usage and can be extremely energy efficient when coded on neur…

Cited by 42SourceScholar
2020

EVDodgeNet: Deep Dynamic Obstacle Dodging with Event Cameras

ICRA 2020poster

Dynamic obstacle avoidance on quadrotors requires low latency. A class of sensors that are particularly suitable for such scenarios are event cameras. In this paper, we present a deep learning based solution for dodging multiple dynamic obstacles on a quadrotor with a single event camera and on-boar…

Cited by 96SourcecodeScholar
2020

Learning Visual Motion Segmentation Using Event Surfaces

CVPR 2020poster

Event-based cameras have been designed for scene motion perception - their high temporal resolution and spatial data sparsity converts the scene into a volume of boundary trajectories and allows to track and analyze the evolution of the scene in time. Analyzing this data is computationally expensive…

Cited by 85PDFScholar
2020

Unsupervised Learning of Dense Optical Flow, Depth and Egomotion with Event-Based Sensors

IROS 2020poster

We present an unsupervised learning pipeline for dense depth, optical flow and egomotion estimation for autonomous driving applications, using the event-based output of the Dynamic Vision Sensor (DVS) as input. The backbone of our pipeline is a bioinspired encoder-decoder neural network architecture…

Cited by 73SourceScholar
2019

EV-IMO: Motion Segmentation Dataset and Learning Pipeline for Event Cameras

IROS 2019poster

We present the first event-based learning approach for motion segmentation in indoor scenes and the first event-based dataset - EV-IMO- which includes accurate pixel-wise motion masks, egomotion and ground truth depth. Our approach is based on an efficient implementation of the SfM learning pipeline…

Cited by 121SourceScholar
2018

Event-Based Moving Object Detection and Tracking

IROS 2018poster

Event-based vision sensors, such as the Dynamic Vision Sensor (DVS), are ideally suited for real-time motion analysis. The unique properties encompassed in the readings of such sensors provide high temporal resolution, superior sensitivity to light and low latency. These properties provide the groun…

Cited by 406SourceScholar
2018

GapFlyt: Active Vision Based Minimalist Structure-Less Gap Detection For Quadrotor Flight

RA-L 2018

Although quadrotors, and aerial robots in general, are inherently active agents, their perceptual capabilities in literature so far have been mostly passive in nature. Researchers and practitioners today use traditional computer vision algorithms with the aim of building a representation of general

Cited by 88SourcecodeScholar
2018

Seeing Behind the Scene: Using Symmetry to Reason About Objects in Cluttered Environments

IROS 2018poster

Symmetry is a common property shared by the majority of man-made objects. This paper presents a novel bottom-up approach for segmenting symmetric objects and recovering their symmetries from 3D pointclouds of natural scenes. Candidate rotational and reflectional symmetries are detected by fitting sy…

Cited by 13SourceScholar
2017

What can i do around here? Deep functional scene understanding for cognitive robots

ICRA 2017poster

For robots that have the capability to interact with the physical environment through their end effectors, understanding the surrounding scenes is not merely a task of image classification or object recognition. To perform actual tasks, it is critical for the robot to have a functional understanding…

Cited by 59SourceScholar
2015

Affordance detection of tool parts from geometric features

ICRA 2015poster

As robots begin to collaborate with humans in everyday workspaces, they will need to understand the functions of tools and their parts. To cut an apple or hammer a nail, robots need to not just know the tool's name, but they must localize the tool's parts and identify their functions. Intuitively, t…

Cited by 386SourceScholar
2015

Contour Detection and Characterization for Asynchronous Event Sensors

ICCV 2015poster

The bio-inspired, asynchronous event-based dynamic vision sensor records temporal changes in the luminance of the scene at high temporal resolution. Since events are only triggered at significant luminance changes, most events occur at the boundary of objects and their parts. The detection of these…

Cited by 29PDFScholar
2015

Detection and Segmentation of 2D Curved Reflection Symmetric Structures

ICCV 2015poster

Symmetry, as one of the key components of Gestalt theory, provides an important mid-level cue that serves as input to higher visual processes such as segmentation. In this work, we propose a complete approach that links the detection of curved reflection symmetries to produce symmetry-constrained se…

Cited by 48PDFScholar
2015

Grasp Type Revisited: A Modern Perspective on a Classical Feature for Vision

CVPR 2015poster

The grasp type provides crucial information about human action. However, recognizing the grasp type in unconstrained scenes is challenging because of the large variations in appearance, occlusions and geometric distortions. In this paper, first we present a convolutional neural network to classify…

Cited by 103SourcePDFScholar
2015

Learning the spatial semantics of manipulation actions through preposition grounding

ICRA 2015poster

In this paper, we introduce an abstract representation for manipulation actions that is based on the evolution of the spatial relations between involved objects. Object tracking in RGBD streams enables straightforward and intuitive ways to model spatial relations in 3D space. Reasoning in 3D overcom…

Cited by 64SourceScholar