← Search

Milan Ganai

4 accepted papers

2026

Self-Supervised Bootstrapping of Action-Predictive Embodied Reasoning

RSS 2026poster

Embodied Chain-of-Thought (CoT) reasoning has significantly enhanced Vision-Language-Action (VLA) models, yet current methods rely on rigid templates to specify reasoning primitives (e.g., objects in the scene, high-level plans, structural affordances). These templates can force policies to process …

Cited by 0SourceScholar
2025

Real-Time Out-of-Distribution Failure Prevention via Multi-Modal Reasoning

CoRL 2025oral

Foundation models can provide robust high-level reasoning on appropriate safety interventions in hazardous scenarios beyond a robot's training data, i.e. out-of-distribution (OOD) failures. However, due to the high inference latency of Large Vision and Language Models, current methods rely on manual…

Cited by 0SourcecodeScholar
2023

Iterative Reachability Estimation for Safe Reinforcement Learning

NeurIPS 2023poster

Ensuring safety is important for the practical deployment of reinforcement learning (RL). Various challenges must be addressed, such as handling stochasticity in the environments, providing rigorous guarantees of persistent state-wise safety satisfaction, and avoiding overly conservative behaviors t…

Cited by 20SourcePDFScholar
2023

Learning Stabilization Control from Observations by Learning Lyapunov-like Proxy Models

ICRA 2023poster

The deployment of Reinforcement Learning to robotics applications faces the difficulty of reward engineering. Therefore, approaches have focused on creating reward functions by Learning from Observations (LfO) which is the task of learning policies from expert trajectories that only contain state se…

Cited by 7SourceScholar