← Search

Kazumi Kasaura

6 accepted papers

2025

Near-Optimal Policy Identification in Robust Constrained Markov Decision Processes via Epigraph Form

ICLR 2025poster

Designing a safe policy for uncertain environments is crucial in real-world control systems. However, this challenge remains inadequately addressed within the Markov decision process (MDP) framework. This paper presents the first algorithm guaranteed to identify a near-optimal policy in a robust con…

2025

Provably Efficient RL under Episode-Wise Safety in Constrained MDPs with Linear Function Approximation

NeurIPS 2025spotlight

We study the reinforcement learning (RL) problem in a constrained Markov decision process (CMDP), where an agent explores the environment to maximize the expected cumulative reward while satisfying a single constraint on the expected total utility value in every episode. While this problem is well u…

Cited by 0SourceScholar
2023

Benchmarking Actor-Critic Deep Reinforcement Learning Algorithms for Robotics Control With Action Constraints

RA-L 2023

This study presents a benchmark for evaluating action-constrained reinforcement learning (RL) algorithms. In action-constrained RL, each action taken by the learning system must comply with certain constraints. These constraints are crucial for ensuring the feasibility and safety of actions in real-

Cited by 22SourcecodeScholar
2022

Prioritized Safe Interval Path Planning for Multi-Agent Pathfinding With Continuous Time on 2D Roadmaps

RA-L 2022

We address a challenging multi-agent pathfinding (MAPF) problem for hundreds of agents moving on a 2D roadmap with continuous time. Despite its known potential for producing better solutions compared to typical grid and discrete-time cases, few approaches have been established to solve this problem

Cited by 29SourceScholar
2022

Uncertainty-Aware Manipulation Planning Using Gravity and Environment Geometry

RA-L 2022

Factory automation robot systems often depend on specially-made jigs that precisely position each part, which increases the system's cost and limits flexibility. We propose a method to determine the 3D pose of an object with high precision and confidence, using only parallel robotic grippers and no

Cited by 9SourceScholar