← Search

Haruki Nishimura

14 accepted papers

2026

A Systematic Study of Data Modalities and Strategies for Co-training Large Behavior Models for Robot Manipulation

RSS 2026poster

Large behavior models (LBMs) have shown strong dexterous manipulation capabilities by extending imitation learning to large-scale training on extensive multi-task robot data, yet their generalization remains limited by the insufficient coverage of available robot data. To expand this coverage withou…

Cited by 0SourceScholar
2026

Beyond Binary Success: Sample-Efficient and Statistically Rigorous Robot Policy Comparison

RSS 2026poster

Generalist robot manipulation policies are becoming increasingly capable, but are limited in evaluation to a small number of hardware rollouts. This strong resource constraint in real-world testing necessitates both more informative performance measures and reliable and efficient evaluation procedur…

Cited by 0SourceScholar
2026

Impact of Different Failures on a Robot’s Perceived Reliability

ICRA 2026poster

Robots fail, potentially leading to a loss in the robot’s perceived reliability (PR), a measure correlated with trustworthiness. In this study we examine how various kinds of failures affect the PR of the robot differently, and how this measure recovers without explicit social repair actions by the …

2025

CUPID: Curating Data your Robot Loves with Influence Functions

CoRL 2025poster

In robot imitation learning, policy performance is tightly coupled with the quality and composition of the demonstration data. Yet, developing a precise understanding of how individual demonstrations contribute to downstream outcomes—such as closed-loop task success or failure—remains a persistent c…

Cited by 0SourceScholar
2025

Can We Detect Failures Without Failure Data? Uncertainty-Aware Runtime Failure Detection for Imitation Learning Policies

RSS 2025poster

Recent years have witnessed impressive robotic manipulation systems driven by advances in imitation learning and generative modeling, such as diffusion- and flow-based approaches. As robot policy performance increases, so does the complexity and time horizon of achievable tasks, inducing unexpected…

Cited by 1PDFScholar
2025

Is Your Imitation Learning Policy Better than Mine? Policy Comparison with Near-Optimal Stopping

RSS 2025poster

Imitation learning has enabled robots to perform complex, long-horizon tasks in challenging dexterous manipulation settings. As new methods are developed, they must be rigorously evaluated and compared against corresponding baselines through repeated evaluation trials. However, policy comparison is…

Cited by 1PDFScholar
2025

SAFE: Multitask Failure Detection for Vision-Language-Action Models

NeurIPS 2025poster

While vision-language-action models (VLAs) have shown promising robotic behaviors across a diverse set of manipulation tasks, they achieve limited success rates when deployed on novel tasks out of the box. To allow these policies to safely interact with their environments, we need a failure detector…

Cited by 0SourcecodeScholar
2025

STITCH-OPE: Trajectory Stitching with Guided Diffusion for Off-Policy Evaluation

NeurIPS 2025spotlight

Off-policy evaluation (OPE) estimates the performance of a target policy using offline data collected from a behavior policy, and is crucial in domains such as robotics or healthcare where direct interaction with the environment is costly or unsafe. Existing OPE methods are ineffective for high-dime…

Cited by 0SourceScholar
2024

How Generalizable is My Behavior Cloning Policy? A Statistical Approach to Trustworthy Performance Evaluation

RA-L 2024

With the rise of stochastic generative models in robot policy learning, end-to-end visuomotor policies are increasingly successful at solving complex tasks by learning from human demonstrations. Nevertheless, since real-world evaluation costs afford users only a small number of policy rollouts, it r

Cited by 15SourcecodeScholar
2023

Residual Q-Learning: Offline and Online Policy Customization without Value

NeurIPS 2023poster

Imitation Learning (IL) is a widely used framework for learning imitative behavior from demonstrations. It is especially appealing for solving complex real-world tasks where handcrafting reward function is difficult, or when the goal is to mimic human expert behavior. However, the learned imitative…

Cited by 6SourcePDFScholar
2022

RAP: Risk-Aware Prediction for Robust Planning

CoRL 2022oral

Robust planning in interactive scenarios requires predicting the uncertain future to make risk-aware decisions. Unfortunately, due to long-tail safety-critical events, the risk is often under-estimated by finite-sampling approximations of probabilistic motion forecasts. This can lead to overconfiden…

Cited by 17SourcecodeScholar
2021

RAT iLQR: A Risk Auto-Tuning Controller to Optimally Account for Stochastic Model Mismatch

RA-L 2021

Successful robotic operation stochastic environments relies on accurate characterization of the underlying probability distributions, yet this is often imperfect due to limited knowledge. This work presents a control algorithm that is capable of handling such distributional mismatches. Specifically,

Cited by 16SourcecodeScholar
2020

Risk-Sensitive Sequential Action Control with Multi-Modal Human Trajectory Forecasting for Safe Crowd-Robot Interaction

IROS 2020poster

This paper presents a novel online framework for safe crowd-robot interaction based on risk-sensitive stochastic optimal control, wherein the risk is modeled by the entropic risk measure. The sampling-based model predictive control relies on mode insertion gradient optimization for this risk measure…

Cited by 49SourceScholar