← Search

Takamitsu Matsubara

44 accepted papers

2026

DAPPER: Discriminability-Aware Policy-To-Policy Preference-Based Reinforcement Learning for Query-Efficient Robot Skill Acquisition

ICRA 2026poster

Preference-based Reinforcement Learning (PbRL) enables policy learning through simple queries comparing trajectories from a single policy, yet suffers from low query efficiency as policy bias limits trajectory diversity and reduces discriminable queries for learning human preferences. This paper ide…

2026

Progressive-Resolution Policy Distillation: Leveraging Coarse-Resolution Simulations for Time-Efficient Fine-Resolution Policy Learning (I)

ICRA 2026poster

In earthwork and construction, excavators often encounter large rocks mixed with various soil conditions, requiring skilled operators. This paper presents a framework for achieving autonomous excavation using reinforcement learning (RL) through a rock excavation simulator. In the simulation, resolut…

Cited by 0Scholar
2026

Tracing Energy Flow: Learning Tactile-Based Grasping Force Control to Reduce Slippage in Dynamic Object Interaction

RA-L 2026

Regulating grasping force to reduce slippage during dynamic object interaction remains a fundamental challenge in robotic manipulation, especially when objects are manipulated by multiple rolling contacts, have unknown properties (such as mass or surface conditions), and when external sensing is unr

Cited by 1SourceScholar
2026

Tracing Energy Flow: Learning Tactile-Based Grasping Force Control to Reduce Slippage in Dynamic Object Interaction

ICRA 2026poster

Regulating grasping force to reduce slippage during dynamic object interaction remains a fundamental challenge in robotic manipulation, especially when objects are manipulated by multiple rolling contacts, have unknown properties (such as mass or surface conditions), and when external sensing is unr…

Cited by 0SourceScholar
2025

Cutting Sequence Diffuser: Sim-to-Real Transferable Planning for Object Shaping by Grinding

RA-L 2025

Automating object shaping by grinding with a robot is a crucial industrial process that involves removing material with a rotating grinding belt. This process generates removal resistance depending on such process conditions as material type, removal volume, and robot grinding posture, all of which

Cited by 0SourceScholar
2025

Feasibility-Aware Imitation Learning from Observations Through a Hand-Mounted Demonstration Interface

ICRA 2025

Imitation learning through a demonstration interface is expected to learn policies for robot automation from intuitive human demonstrations. However, due to the differences in human and robot movement characteristics, a human expert might unintentionally demonstrate an action that the robot cannot e

Cited by 1SourceScholar
2025

ICCO: Learning an Instruction-conditioned Coordinator for Language-guided Task-aligned Multi-robot Control

IROS 2025

Recent advances in Large Language Models (LLMs) have permitted the development of language-guided multi-robot systems, which allow robots to execute tasks based on natural language instructions. However, achieving effective coordination in distributed multi-agent environments remains challenging due

Cited by 1SourcecodeScholar
2025

Reinforcement Learning of Flexible Policies for Symbolic Instructions With Adjustable Mapping Specifications

RA-L 2025

Symbolic task representation is a powerful tool for encoding human instructions and domain knowledge. Such instructions guide robots to accomplish diverse objectives and meet constraints through reinforcement learning (RL). Most existing methods are based on fixed mappings from environmental states

Cited by 2SourceScholar
2024

Domain Randomization-free Sim-to-Real : An Attention-Augmented Memory Approach for Robotic Tasks

IROS 2024poster

The sim-to-real gap, a long-standing challenge in the field of robotics, has garnered significant attention. Essentially, it is important to learn robust representation models that can be seamlessly applied in both simulation and real world. Traditional approaches like domain randomization have demo…

Cited by 0SourceScholar
2024

Leveraging Demonstrator-Perceived Precision for Safe Interactive Imitation Learning of Clearance-Limited Tasks

RA-L 2024

Interactive imitation learning is an efficient, model-free method through which a robot can learn a task by repetitively iterating an execution of a learning policy and a data collection by querying human demonstrations. However, deploying unmatured policies for clearance-limited tasks, like industr

Cited by 6SourceScholar
2023

Disturbance Injection Under Partial Automation: Robust Imitation Learning for Long-Horizon Tasks

RA-L 2023

Partial Automation (PA) with intelligent support systems has been introduced in industrial machinery and advanced automobiles to reduce the burden of long hours of human operation. Under PA, operators perform manual operations (providing actions) and operations that switch to automatic/manual mode (

Cited by 5SourceScholar
2023

Domains as Objectives: Multi-Domain Reinforcement Learning with Convex-Coverage Set Learning for Domain Uncertainty Awareness

IROS 2023poster

Domain randomization (DR) is a powerful framework that has allowed the transfer of policies from randomized domain (a.k.a. simulation) to real robots with little to no retraining requirement. However, because the policy has to perform well for many different domain conditions, DR tends to produce su…

Cited by 1SourceScholar
2023

Learning to Shape by Grinding: Cutting-Surface-Aware Model-Based Reinforcement Learning

RA-L 2023

Object shaping by grinding is a crucial industrial process in which a rotating grinding belt removes material. Object-shape transition models are essential to achieving automation by robots; however, learning such a complex model that depends on process conditions is challenging because it requires

Cited by 9SourceScholar
2023

Reinforcement Learning With Energy-Exchange Dynamics for Spring-Loaded Biped Robot Walking

RA-L 2023

This paper presents a probabilistic Model-based Reinforcement Learning (MBRL) approach for learning the Energy-exchange Dynamics (EED) of a spring-loaded biped robot. Our approach enables on-site walking acquisition with high sample efficiency, real-time planning capability, and generalizability acr

Cited by 7SourceScholar
2023

Reinforcement Learning of Action and Query Policies With LTL Instructions Under Uncertain Event Detector

RA-L 2023

Reinforcement learning (RL) with linear temporal logic (LTL) objectives can allow robots to carry out symbolic event plans in unknown environments. Most existing methods assume that the event detector can accurately map environmental states to symbolic events; however, uncertainty is inevitable for

Cited by 7SourceScholar
2022

Disturbance-injected Robust Imitation Learning with Task Achievement

ICRA 2022poster

Robust imitation learning using disturbance injections overcomes issues of limited variation in demonstrations. However, these methods assume demonstrations are optimal, and that policy stabilization can be learned via simple augmentations. In real-world scenarios, demonstrations are often of divers…

Cited by 13SourceScholar
2022

Gaussian Process Self-triggered Policy Search in Weakly Observable Environments

ICRA 2022poster

The environments of such large industrial machines as waste cranes in waste incineration plants are often weakly observable, where little information about the environ-mental state is contained in the observations due to technical difficulty or maintenance cost (e.g., no sensors for observing the st…

Cited by 3SourceScholar
2022

Physically Consistent Preferential Bayesian Optimization for Food Arrangement

RA-L 2022

This letter considers the problem of estimating a preferred food arrangement for users from interactive pairwise comparisons using Computer Graphics (CG)-based dish images. As a foodservice industry requirement, we need to utilize domain rules for the geometry of the arrangement (e.g., the food layo

Cited by 1SourceScholar
2022

Randomized-to-Canonical Model Predictive Control for Real-World Visual Robotic Manipulation

RA-L 2022

Many works have recently explored Sim-to-real transferable visual model predictive control (MPC). However, such works are limited to one-shot transfer, where real-world data must be collected once to perform the sim-to-real transfer, which remains a significant human effort in transferring the model

Cited by 5SourceScholar
2022

Uncertainty-Aware Manipulation Planning Using Gravity and Environment Geometry

RA-L 2022

Factory automation robot systems often depend on specially-made jigs that precisely position each part, which increases the system's cost and limits flexibility. We propose a method to determine the 3D pose of an object with high precision and confidence, using only parallel robotic grippers and no

Cited by 9SourceScholar
2021

Bayesian Disturbance Injection: Robust Imitation Learning of Flexible Policies

ICRA 2021poster

Scenarios requiring humans to choose from multiple seemingly optimal actions are commonplace, however standard imitation learning often fails to capture this behavior. Instead, an over-reliance on replicating expert actions induces inflexible and unstable policies, leading to poor generalizability i…

Cited by 10SourceScholar
2021

Binarized P-Network: Deep Reinforcement Learning of Robot Control from Raw Images on FPGA

RA-L 2021

This letter explores a deep reinforcement learning (DRL) approach for designing image-based control for edge robots to be implemented on Field Programmable Gate Arrays (FPGAs). Although FPGAs are more power-efficient than CPUs and GPUs, a typical DRL method cannot be applied since they are composed

Cited by 9SourceScholar
2021

Deep reinforcement learning of event-triggered communication and control for multi-agent cooperative transport

ICRA 2021poster

In this paper, we explore a multi-agent reinforcement learning approach to address the design problem of communication and control strategies for multi-agent cooperative transport. Typical end-to-end deep neural network policies may be insufficient for covering communication and control; these metho…

Cited by 23SourceScholar
2021

Learning Robotic Contact Juggling

IROS 2021poster

Robotic contact juggling is a challenging task in which robots must control the movement of a ball rapidly and indirectly without holding it while keeping the ball in and sometimes out of contact with the robot’s body. In this work, we address the problem of learning such robotic contact juggling fr…

Cited by 4SourceScholar
2021

Uncertainty-Aware Contact-Safe Model-Based Reinforcement Learning

RA-L 2021

This letter presents contact-safe Model-based Reinforcement Learning (MBRL) for robot applications that achieves contact-safe behaviors in the learning process. In typical MBRL, we cannot expect the data-driven model to generate accurate and reliable policies to the intended robotic tasks during the

Cited by 21SourceScholar
2020

Bayesian Policy Optimization for Waste Crane With Garbage Inhomogeneity

RA-L 2020

The objective of this study is to develop a framework that can optimize control policies of a waste crane at a waste incineration plant through an autonomous trial and error manner. Since a waste crane is a massive mechanical system that moves slowly and takes several minutes to execute a task, obta

Cited by 7SourceScholar
2020

Contact-based in-hand pose estimation using Bayesian state estimation and particle filtering

ICRA 2020poster

In industrial assembly tasks, the position of an object grasped by the robot has to be known with high precision in order to insert or place it. In real applications, this problem is commonly solved by jigs that are specially produced for each part. However, they significantly limit flexibility and…

Cited by 30SourceScholar
2020

Dynamic Actor-Advisor Programming for Scalable Safe Reinforcement Learning

ICRA 2020poster

Real-world robots have complex strict constraints. Therefore, safe reinforcement learning algorithms that can simultaneously minimize the total cost and the risk of constraint violation are crucial. However, almost no algorithms exist that can scale to high-dimensional systems to the best of our kno…

Cited by 8SourceScholar
2020

Exploiting Visual-Outer Shape for Tactile-Inner Shape Estimation of Objects Covered with Soft Materials

RA-L 2020

In this letter, we consider the problem of inner-shape estimation of objects covered with soft materials, e.g., pastries wrapped in paper or vinyl, water bottles covered with shock-absorbing fabrics, or human bodies dressed in clothes. Due to the softness of the covered materials, tactile informatio

Cited by 2SourceScholar
2020

Learning Force Control for Contact-Rich Manipulation Tasks With Rigid Position-Controlled Robots

RA-L 2020

Reinforcement Learning (RL) methods have been proven successful in solving manipulation tasks autonomously. However, RL is still not widely adopted on real robotic systems because working with real hardware entails additional challenges, especially when using rigid position-controlled manipulators.

Cited by 137SourceScholar
2020

Learning Soft Robotic Assembly Strategies from Successful and Failed Demonstrations

IROS 2020poster

Physically soft robots are promising for robotic assembly tasks as they allow stable contacts with the environment. In this study, we propose a novel learning system for soft robotic assembly strategies. We formulate this problem as a reinforcement learning task and design the reward function from h…

Cited by 24SourceScholar
2020

Quaternion-Based Trajectory Optimization of Human Postures for Inducing Target Muscle Activation Patterns

RA-L 2020

In exercise and rehabilitation, to effectively train the human body, human motion trajectory is essential because it induces muscle activity patterns. In this letter, we develop a novel framework for the trajectory optimization of human postures, including the head, the limbs, and the body to induce

Cited by 6SourceScholar
2020

Sample-and-computation-efficient Probabilistic Model Predictive Control with Random Features

ICRA 2020poster

Gaussian processes (GPs) based Reinforcement Learning (RL) methods with Model Predictive Control (MPC) have demonstrated their excellent sample efficiency. However, since the computational cost of GPs largely depends on the training sample size, learning an accurate dynamics using GPs result in low…

Cited by 10SourceScholar
2019

Exploiting Human and Robot Muscle Synergies for Human-in-the-loop Optimization of EMG-based Assistive Strategies

ICRA 2019poster

In this study, we propose a novel human-in-the-loop optimization approach for exoskeleton robot control. We develop a method to optimize widely-used Electromyography (EMG)-based assistive strategies. If we use multiple EMG channels to control multi-DoF robots, optimization process becomes complex an…

Cited by 12SourceScholar
2019

Probabilistic Active Filtering for Object Search in Clutter

ICRA 2019poster

This paper proposes a probabilistic approach for object search in clutter. Due to heavy occlusions, it is vital for an agent to be able to gradually reduce uncertainty in observations of the objects in its workspace by systematically rearranging them. Probabilistic methodologies present a promising…

Cited by 9SourceScholar
2019

Reinforcement Learning Boat Autopilot: A Sample-efficient and Model Predictive Control based Approach

IROS 2019poster

In this research we focus on developing a reinforcement learning system for a challenging task: autonomous control of a real-sized boat, with difficulties arising from large uncertainties in the challenging ocean environment and the extremely high cost of exploring and sampling with a real boat. To…

Cited by 41SourceScholar
2017

Deep dynamic policy programming for robot control with raw images

IROS 2017poster

Deep reinforcement learning has drawn much attention in robot control since it enables agents to learn control policies from very high dimensional states such as raw images. On the other hand, its dependency upon the availability of a significant quantity of training samples and its fragility in lea…

Cited by 16SourceScholar
2017

Learning task-parametrized assistive strategies for exoskeleton robots by multi-task reinforcement learning

ICRA 2017poster

Recent studies suggest that reinforcement learning has great potential for generating assistive strategies in exoskeletons through physical interactions between a user and a robot. Previous methods focused on a task-specific assistive strategy, where for every single task (situation/context), the us…

Cited by 24SourceScholar
2017

Local driving assistance from demonstration for mobility aids

ICRA 2017poster

Active assistive mobility systems are largely limited to a-priori mapped environments, whereas their reactive assistive counterparts are in general location independent and focus on the provision of collision avoidance in the immediate space surrounding the platform. This paper presents a framework…

Cited by 8SourceScholar
2017

User-robot collaborative excitation for PAM model identification in exoskeleton robots

IROS 2017poster

Pneumatic Artificial Muscle (PAM) actuators have been used as exoskeletons because of their inherited compliance and high power-weight ratio. However, creating accurate models remains difficult mainly due to the compliance issue; the model can be changed by the force applied by the user. Therefore,…

Cited by 11SourceScholar
2016

Learning assistive strategies from a few user-robot interactions: Model-based reinforcement learning approach

ICRA 2016

Designing an assistive strategy for exoskeletons is a key ingredient in movement assistance and rehabilitation. While several approaches have been explored, most studies are based on mechanical models of the human user, i.e., rigid-body dynamics or Center of Mass (CoM)-Zero Moment Point (ZMP) invert

Cited by 25SourceScholar