← Search

Florian Shkurti

45 accepted papers

2026

SICNav-Diffusion: Safe and Interactive Crowd Navigation with Diffusion Trajectory Predictions

ICRA 2026poster

To navigate crowds without collisions, robots must interact with humans by forecasting their future motion and reacting accordingly. While learning-based prediction models have shown success in generating likely human trajectory predictions, integrating these stochastic models into a robot controlle…

2026

SICNav: Safe and Interactive Crowd Navigation Using Model Predictive Control and Bilevel Optimization (Abstract Reprint)

AAAI 2026technical

Robots need to predict and react to human motions to navigate through a crowd without collisions. Many existing methods decouple prediction from planning, which does not account for the interaction between robot and human motions and can lead to the robot getting stuck. We propose SICNav, a Model Pr

Cited by 0SourcePDFScholar
2025

AnyPlace: Learning Generalizable Object Placement for Robot Manipulation

CoRL 2025poster

Object placement in robotic tasks is inherently challenging due to the diversity of object geometries and placement configurations. We address this with AnyPlace, a two-stage method trained entirely on synthetic data, capable of predicting a wide range of feasible placement poses for real-world task…

Cited by 0SourcecodeScholar
2025

Automated Planning Domain Inference for Task and Motion Planning

ICRA 2025

Task and motion planning (TAMP) frameworks address long and complex planning problems by integrating high-level task planners with low-level motion planners. However, existing TAMP methods rely heavily on the manual design of planning domains that specify the preconditions and postconditions of all

Cited by 5SourceScholar
2025

Gaussian Splatting Visual MPC for Granular Media Manipulation

ICRA 2025

Recent advancements in learned 3D representations have enabled significant progress in solving complex robotic manipulation tasks, particularly for rigid-body objects. However, manipulating granular materials such as beans, nuts, and rice remains challenging due to the intricate physics of particle

Cited by 0SourceScholar
2025

Generalization in VAE and Diffusion Models: A Unified Information-Theoretic Analysis

ICLR 2025poster

Despite the empirical success of Diffusion Models (DMs) and Variational Autoencoders (VAEs), their generalization performance remains theoretically underexplored, especially lacking a full consideration of the shared encoder-generator structure. Leveraging recent information-theoretic tools, we prop…

Cited by 0SourcePDFScholar
2025

Model AI Assignments 2025

AAAI 2025technical

The Model AI Assignments session seeks to gather and disseminate the best assignment designs of the Artificial Intelligence (AI) Education community. Recognizing that assignments form the core of student learning experience, we here present abstracts of thirteen AI assignments from the 2025 session…

Cited by 0SourcePDFScholar
2025

SAFE: Multitask Failure Detection for Vision-Language-Action Models

NeurIPS 2025poster

While vision-language-action models (VLAs) have shown promising robotic behaviors across a diverse set of manipulation tasks, they achieve limited success rates when deployed on novel tasks out of the box. To allow these policies to safely interact with their environments, we need a failure detector…

Cited by 0SourcecodeScholar
2025

SICNav-Diffusion: Safe and Interactive Crowd Navigation With Diffusion Trajectory Predictions

RA-L 2025

To navigate crowds without collisions, robots must interact with humans by forecasting their future motion and reacting accordingly. While learning-based prediction models have shown success in generating likely human trajectory predictions, integrating these stochastic models into a robot controlle

Cited by 10SourcecodeScholar
2025

STAMP: Differentiable Task and Motion Planning via Stein Variational Gradient Descent

RA-L 2025

Planning for sequential robotics tasks often requires integrated symbolic and geometric reasoning. TAMP algorithms typically solve these problems by performing a tree search over high-level task sequences while checking for kinematic and dynamic feasibility. This can be inefficient because, typicall

Cited by 8SourceScholar
2025

STITCH-OPE: Trajectory Stitching with Guided Diffusion for Off-Policy Evaluation

NeurIPS 2025spotlight

Off-policy evaluation (OPE) estimates the performance of a target policy using offline data collected from a behavior policy, and is crucial in domains such as robotics or healthcare where direct interaction with the environment is costly or unsafe. Existing OPE methods are ineffective for high-dime…

Cited by 0SourceScholar
2025

Search-TTA: A Multi-Modal Test-Time Adaptation Framework for Visual Search in the Wild

CoRL 2025poster

To perform autonomous visual search for environmental monitoring, a robot may leverage satellite imagery as a prior map. This can help inform coarse, high level search and exploration strategies, even when such images lack sufficient resolution to allow fine-grained, explicit visual recognition of t…

Cited by 0SourceScholar
2025

Synthetica: Large Scale Synthetic Data Generation for Robot Perception

IROS 2025

Vision-based object detectors are a crucial basis for robotics applications as they provide valuable information about object localization in the environment. These need to ensure high reliability in different lighting conditions, occlusions, and visual artifacts, all while running in real-time. Col

Cited by 6SourceScholar
2024

ConceptGraphs: Open-Vocabulary 3D Scene Graphs for Perception and Planning

ICRA 2024poster

For robots to perform a wide variety of tasks, they require a 3D representation of the world that is semantically rich, yet compact and efficient for task-driven perception and planning. Recent approaches have attempted to leverage features from large vision-language models to encode semantics in 3D…

Cited by 202SourceScholar
2023

ConceptFusion: Open-set multimodal 3D mapping

RSS 2023poster

Building 3D maps of the environment is central to robot navigation, planning, and interaction with objects in a scene. Most existing approaches that integrate semantic concepts with 3D maps largely remain confined to the closed-set setting: they can only reason about a finite set of concepts, pre-de…

2023

Generating Transferable Adversarial Simulation Scenarios for Self-Driving via Neural Rendering

CoRL 2023poster

Self-driving software pipelines include components that are learned from a significant number of training examples, yet it remains challenging to evaluate the overall system's safety and generalization performance. Together with scaling up the real-world deployment of autonomous vehicles, it is of c…

Cited by 3SourceScholar
2023

MVTrans: Multi-View Perception of Transparent Objects

ICRA 2023poster

Transparent object perception is a crucial skill for applications such as robot manipulation in household and laboratory settings. Existing methods utilize RGB-D or stereo inputs to handle a subset of perception tasks including depth and pose estimation. However transparent object perception remains…

Cited by 26SourcecodeScholar
2023

Policy-Guided Lazy Search with Feedback for Task and Motion Planning

ICRA 2023poster

PDDLStream solvers have recently emerged as viable solutions for Task and Motion Planning (TAMP) problems, extending PDDL to problems with continuous action spaces. Prior work has shown how PDDLStream problems can be reduced to a sequence of PDDL planning problems, which can then be solved using off…

Cited by 11SourceScholar
2023

Preserving Linear Separability in Continual Learning by Backward Feature Projection

CVPR 2023poster

Catastrophic forgetting has been a major challenge in continual learning, where the model needs to learn new tasks with limited or no access to data from previously seen tasks. To tackle this challenge, methods based on knowledge distillation in feature space have been proposed and shown to reduce f…

2023

Sparsifiner: Learning Sparse Instance-Dependent Attention for Efficient Vision Transformers

CVPR 2023poster

Vision Transformers (ViT) have shown competitive advantages in terms of performance compared to convolutional neural networks (CNNs), though they often come with high computational costs. To this end, previous methods explore different attention patterns by limiting a fixed number of spatially nearb…

Cited by 19SourcePDFScholar
2023

Stochastic Planning for ASV Navigation Using Satellite Images

ICRA 2023poster

Autonomous surface vessels (ASV) represent a promising technology to automate water-quality monitoring of lakes. In this work, we use satellite images as a coarse map and plan sampling routes for the robot. However, inconsistency between the satellite images and the actual lake, as well as environme…

Cited by 8SourceScholar
2022

SLIC: Self-Supervised Learning With Iterative Clustering for Human Action Videos

CVPR 2022oral

Self-supervised methods have significantly closed the gap with end-to-end supervised learning for image classification [13,24]. In the case of human action videos, however, where both appearance and motion are significant factors of variation, this gap remains significant [28,58]. One of the key rea…

Cited by 37PDFcodeScholar
2021

Conservative Safety Critics for Exploration

ICLR 2021poster

Safe exploration presents a major challenge in reinforcement learning (RL): when active data collection requires deploying partially trained policies, we must ensure that these policies avoid catastrophically unsafe regions, while still enabling trial and error learning. In this paper, we target the…

Cited by 168SourcePDFScholar
2021

Continual Model-Based Reinforcement Learning with Hypernetworks

ICRA 2021poster

Effective planning in model-based reinforcement learning (MBRL) and model-predictive control (MPC) relies on the accuracy of the learned dynamics model. In many instances of MBRL and MPC, this model is assumed to be stationary and is periodically re-trained from scratch on state transition experienc…

Cited by 61SourceScholar
2021

DIBS: Diversity Inducing Information Bottleneck in Model Ensembles

AAAI 2021technical

Although deep learning models have achieved state-of-the art performance on a number of vision tasks, generalization over high dimensional multi-modal data, and reliable predictive uncertainty estimation are still active areas of research. Bayesian approaches including Bayesian Neural Nets (BNNs) d…

Cited by 53SourcePDFScholar
2021

LOHO: Latent Optimization of Hairstyles via Orthogonalization

CVPR 2021poster

Hairstyle transfer is challenging due to hair structure differences in the source and target hair. Therefore, we propose Latent Optimization of Hairstyles via Orthogonalization (LOHO), an optimization-based approach using GAN inversion to infill missing hair structure details in latent space during…

Cited by 60PDFcodeScholar
2021

Latent Attention Augmentation for Robust Autonomous Driving Policies

IROS 2021poster

Model-free reinforcement learning has become a viable approach for vision-based robot control. However, sample complexity and adaptability to domain shifts remain persistent challenges when operating in high-dimensional observation spaces (images, LiDAR), such as those that are involved in autonomou…

Cited by 4SourceScholar
2021

Latent Skill Planning for Exploration and Transfer

ICLR 2021poster

To quickly solve new tasks in complex environments, intelligent agents need to build up reusable knowledge. For example, a learned world model captures knowledge about the environment that applies to new tasks. Similarly, skills capture general behaviors that can apply to new tasks. In this paper, w…

Cited by 19SourcePDFScholar
2021

Physics-Based Human Motion Estimation and Synthesis From Videos

ICCV 2021poster

Human motion synthesis is an important problem for applications in graphics and gaming, and even in simulation environments for robotics. Existing methods require accurate motion capture data for training, which is costly to obtain. Instead, we propose a framework for training generative models of p…

Cited by 110PDFScholar
2021

Seeing Glass: Joint Point-Cloud and Depth Completion for Transparent Objects

CoRL 2021oral

The basis of many object manipulation algorithms is RGB-D input. Yet, commodity RGB-D sensors can only provide distorted depth maps for a wide range of transparent objects due light refraction and absorption. To tackle the perception challenges posed by transparent objects, we propose TranspareNet,…

Cited by 62SourceScholar
2021

Shaping Rewards for Reinforcement Learning with Imperfect Demonstrations using Generative Models

ICRA 2021poster

The potential benefits of model-free reinforcement learning to real robotics systems are limited by its uninformed exploration that leads to slow convergence, lack of data-efficiency, and unnecessary interactions with the environment. To address these drawbacks we propose a method that combines rein…

Cited by 35SourceScholar
2021

Taskography: Evaluating robot task planning over large 3D scene graphs

CoRL 2021poster

3D scene graphs (3DSGs) are an emerging description; unifying symbolic, topological, and metric scene representations. However, typical 3DSGs contain hundreds of objects and symbols even for small environments; rendering task planning on the \emph{full} graph impractical. We construct \textbf{Taskog…

Cited by 84SourcecodeScholar
2021

gradSim: Differentiable simulation for system identification and visuomotor control

ICLR 2021poster

In this paper, we tackle the problem of estimating object physical properties such as mass, friction, and elasticity directly from video sequences. Such a system identification problem is fundamentally ill-posed due to the loss of information during image formation. Current best solutions to the pro…

Cited by 40SourcePDFScholar
2020

Catch the Ball: Accurate High-Speed Motions for Mobile Manipulators via Inverse Dynamics Learning

IROS 2020poster

Mobile manipulators consist of a mobile platform equipped with one or more robot arms and are of interest for a wide array of challenging tasks because of their extended workspace and dexterity. Typically, mobile manipulators are deployed in slow-motion collaborative robot scenarios. In this paper,…

Cited by 34SourceScholar
2020

One-Shot Informed Robotic Visual Search in the Wild

IROS 2020poster

We consider the task of underwater robot navigation for the purpose of collecting scientifically relevant video data for environmental monitoring. The majority of field robots that currently perform monitoring tasks in unstructured natural environments navigate via path-tracking a pre-specified sequ…

Cited by 16SourcecodeScholar
2020

Vision-Based Goal-Conditioned Policies for Underwater Navigation in the Presence of Obstacles

RSS 2020poster

We present Nav2Goal, a data-efficient and end-to-end learning method for goal-conditioned visual navigation. Our technique is used to train a navigation policy that enables a robot to navigate close to sparse geographic waypoints provided by a user without any prior map, all while avoiding obstacles…

Cited by 63SourcePDFScholar
2019

Generating Adversarial Driving Scenarios in High-Fidelity Simulators

ICRA 2019poster

In recent years self-driving vehicles have become more commonplace on public roads, with the promise of bringing safety and efficiency to modern transportation systems. Increasing the reliability of these vehicles on the road requires an extensive suite of software tests, ideally performed on high-f…

Cited by 173SourceScholar
2017

Underwater multi-robot convoying using visual tracking by detection

IROS 2017poster

We present a robust multi-robot convoying approach that relies on visual detection of the leading agent, thus enabling target following in unstructured 3-D environments. Our method is based on the idea of tracking-by-detection, which interleaves efficient model-based object detection with temporal f…

Cited by 81SourcecodeScholar