← Search

Michael M. Zavlanos

15 accepted papers

2025

Distributionally Robust Multi-Agent Reinforcement Learning for Dynamic Chute Mapping

ICML 2025poster

In Amazon robotic warehouses, the destination-to-chute mapping problem is crucial for efficient package sorting. Often, however, this problem is complicated by uncertain and dynamic package induction rates, which can lead to increased package recirculation. To tackle this challenge, we introduce a D…

Cited by 0SourcePDFScholar
2025

Enhancing Cooperative Multi-Agent Reinforcement Learning with State Modelling and Adversarial Exploration

ICML 2025poster

Learning to cooperate in distributed partially observable environments with no communication abilities poses significant challenges for multi-agent deep reinforcement learning (MARL). This paper addresses key concerns in this domain, focusing on inferring state representations from individual agent…

2024

Outlier-Robust Distributionally Robust Optimization via Unbalanced Optimal Transport

NeurIPS 2024poster

Distributionally Robust Optimization (DRO) accounts for uncertainty in data distributions by optimizing the model performance against the worst possible distribution within an ambiguity set. In this paper, we propose a DRO framework that relies on a new distance inspired by Unbalanced Optimal Transp…

Cited by 14SourcePDFScholar
2023

Policy Stitching: Learning Transferable Robot Policies

CoRL 2023poster

Training robots with reinforcement learning (RL) typically involves heavy interactions with the environment, and the acquired skills are often sensitive to changes in task environments and robot kinematics. Transfer RL aims to leverage previous knowledge to accelerate learning of new tasks or new bo…

Cited by 9SourceScholar
2022

Formal Verification of Stochastic Systems with ReLU Neural Network Controllers

ICRA 2022poster

In this work, we address the problem of formal safety verification for stochastic cyber-physical systems (CPS) equipped with ReLU neural network (NN) controllers. Our goal is to find the set of initial states from where, with a predetermined confidence, the system will not reach an unsafe configurat…

Cited by 7SourceScholar
2022

Receding Horizon Tracking of an Unknown Number of Mobile Targets using a Bearings-Only Sensor

ICRA 2022poster

Planning the motion of bearings-only sensors is critical for enabling accurate tracking of the positions of moving targets. In this paper, we demonstrate planning the observer's motion over horizons greater than one step for estimating an unknown and varying number of indistinguishable, maneuvering…

Cited by 3SourceScholar
2021

Bearing-Only Active Sensing Under Merged Measurements

RA-L 2021

In this letter we propose an algorithm to actively track multiple moving targets using a bearing-only sensor in the presence of merged measurements. Merged measurements arise from sensor resolution constraints and therefore targets that are close in relative bearing to the sensor get reported as a s

Cited by 8SourceScholar
2021

Distance Estimation Using Self-Induced Noise of an Aerial Vehicle

RA-L 2021

In this letter, we propose an algorithm to estimate the distance between an aerial vehicle and a large obstacle using the self-induced noise present during the vehicle's normal operation. We demonstrate the feasibility of using the proposed estimation method in real-time as a feedback mechanism to a

Cited by 6SourceScholar
2021

Model-Free Reinforcement Learning for Stochastic Games with Linear Temporal Logic Objectives

ICRA 2021poster

We study synthesis of control strategies from linear temporal logic (LTL) objectives in unknown environments. We model this problem as a turn-based zero-sum stochastic game between the controller and the environment, where the transition probabilities and the model topology are fully unknown. The wi…

Cited by 20SourceScholar
2020

Bio-Inspired Distance Estimation using the Self-Induced Acoustic Signature of a Motor-Propeller System

ICRA 2020poster

In this paper we propose an algorithm to actively control the distance of a motor-propeller system (MPS) to a large obstacle using data from a single microphone. The method is based upon a broadband constructive/destructive interference pattern across the audible frequency band that is present when…

Cited by 6SourceScholar
2020

Control Synthesis from Linear Temporal Logic Specifications using Model-Free Reinforcement Learning

ICRA 2020poster

We present a reinforcement learning (RL) frame-work to synthesize a control policy from a given linear temporal logic (LTL) specification in an unknown stochastic environment that can be modeled as a Markov Decision Process (MDP). Specifically, we learn a policy that maximizes the probability of sat…

Cited by 168SourceScholar
2020

Deep Imitative Reinforcement Learning for Temporal Logic Robot Motion Planning with Noisy Semantic Observations

ICRA 2020poster

In this paper, we propose a Deep Imitative Q-learning (DIQL) method to synthesize control policies for mobile robots that need to satisfy Linear Temporal Logic (LTL) specifications using noisy semantic observations of their surroundings. The robot sensing error is modeled using probabilistic labels…

Cited by 9SourceScholar
2018

Distributed Intermittent Communication Control of Mobile Robot Networks Under Time-Critical Dynamic Tasks

ICRA 2018poster

In this paper, we develop a distributed intermittent communication framework for teams of mobile robots that are responsible for accomplishing time-critical dynamic tasks and sharing the collected information with all other robots and possibly also with a user. Specifically, we consider situations w…

Cited by 18SourceScholar