← Search

Takayuki Osa

21 accepted papers

2026

Cross-Embodiment Offline Reinforcement Learning for Heterogeneous Robot Datasets

ICLR 2026poster

Scalable robot policy pre-training has been hindered by the high cost of collecting high-quality demonstrations for each platform. In this study, we address this issue by uniting offline reinforcement learning (offline RL) with cross-embodiment learning. Offline RL leverages both expert and abundant…

Cited by 0SourceScholar
2026

Rethinking Policy Diversity in Ensemble Policy Gradient in Large-Scale Reinforcement Learning

ICLR 2026poster

Scaling reinforcement learning to tens of thousands of parallel environments requires overcoming the limited exploration capacity of a single policy. Ensemble-based policy gradient methods, which employ multiple policies to collect diverse samples, have recently been proposed to promote exploration.…

Cited by 0SourcecodeScholar
2026

Revisiting Regularized Policy Optimization for Stable and Efficient Reinforcement Learning in Two-Player Games

ICML 2026poster

Two-player games such as board games have long been used as traditional benchmark for reinforcement learning. This work revisits a regularized policy optimization with reverse Kullback-Leibler divergence and entropy divergence and analyzes this combination on two-player zero-sum settings from theore…

Cited by 0SourceScholar
2026

Unsupervised Domain Adaptation for Robust Imitation Learning under Visual Perturbations

ICRA 2026poster

Vision-based robot manipulation systems often suffer from performance degradation under domain shifts in visual inputs. While data augmentation is commonly employed in reinforcement learning, its application in imitation learning remains relatively underexplored. Our preliminary experiments indicate…

Cited by 0Scholar
2025

Gradual Transition from Bellman Optimality Operator to Bellman Operator in Online Reinforcement Learning

ICML 2025poster

For continuous action spaces, actor-critic methods are widely used in online reinforcement learning (RL). However, unlike RL algorithms for discrete actions, which generally model the optimal value function using the Bellman optimality operator, RL algorithms for continuous actions typically model Q…

2025

Uncertainty-aware Motion Planning based on Stochastic Forward/Inverse Kinematics Models for Tensegrity Manipulators

IROS 2025

Robots whose shape and stiffness are determined by internal forces generally have complex shape-stiffness relationships that depend on their structure. As a result, there are difficulties such as a decrease in shape reproducibility when the robot is not stiff, and a decrease in the range of motion w

Cited by 0SourceScholar
2024

Active Learning for Forward/Inverse Kinematics of Redundantly-driven Flexible Tensegrity Manipulator

IROS 2024poster

In flexible redundantly-driven multi-DOF systems, like living beings, the representation of redundant kinematics including the diversity of solutions, is crucial for leveraging its distinctive characteristics. This paper proposes an active learning framework for forward and inverse modeling of compl…

Cited by 0SourceScholar
2024

Open X-Embodiment: Robotic Learning Datasets and RT-X Models : Open X-Embodiment Collaboration

ICRA 2024

Large, high-capacity models trained on diverse datasets have shown remarkable successes on efficiently tackling downstream applications. In domains from NLP to Computer Vision, this has led to a consolidation of pretrained models, with general pretrained backbones serving as a starting point for man

Cited by 910SourcecodeScholar
2024

Open X-Embodiment: Robotic Learning Datasets and RT-X Models : Open X-Embodiment Collaboration0

ICRA 2024poster

Large, high-capacity models trained on diverse datasets have shown remarkable successes on efficiently tackling downstream applications. In domains from NLP to Computer Vision, this has led to a consolidation of pretrained models, with general pretrained backbones serving as a starting point for man…

Cited by 259SourcecodeScholar
2024

Robustifying a Policy in Multi-Agent RL with Diverse Cooperative Behaviors and Adversarial Style Sampling for Assistive Tasks

ICRA 2024poster

Autonomous assistance of people with motor impairments is one of the most promising applications of autonomous robotic systems. Recent studies have reported encouraging results using deep reinforcement learning (RL) in the healthcare domain. Previous studies showed that assistive tasks can be formul…

Cited by 1SourceScholar
2024

Symmetric Q-learning: Reducing Skewness of Bellman Error in Online Reinforcement Learning

AAAI 2024technical

In deep reinforcement learning, estimating the value function to evaluate the quality of states and actions is essential. The value function is often trained using the least squares method, which implicitly assumes a Gaussian error distribution. However, a recent study suggested that the error distr…

Cited by 0SourcePDFScholar
2024

Touch-Based Manipulation with Multi-Fingered Robot using Off-policy RL and Temporal Contrastive Learning

ICRA 2024poster

Tactile information holds promise for enhancing the manipulation capabilities of multi-fingered robots. In tasks such as in-hand manipulation, where robots frequently switch between contact and non-contact states, it is important to address the partial observability of tactile sensors and to properl…

Cited by 1SourceScholar
2023

Forward/Inverse Kinematics Modeling for Tensegrity Manipulator Based on Goal-Conditioned Variational Autoencoder

IROS 2023poster

This paper uses a data-driven approach to model a highly redundantly driven tensegrity manipulator's forward and inverse kinematics. The tensegrity manipulator is based on a class-1 tensegrity with 20 struts and bends by 40 pneumatic actuators whose internal pressures are independently controlled. B…

Cited by 4SourceScholar
2023

Learning Adaptive Policies for Autonomous Excavation Under Various Soil Conditions by Adversarial Domain Sampling

RA-L 2023

Excavation is a frequent task in construction. In this context, automation is expected to reduce hazard risks and labor-intensive work. To this end, recent studies have investigated using reinforcement learning (RL) to automate construction machines. One of the challenges in applying RL to excavatio

Cited by 5SourceScholar
2020

Hierarchical Stochastic Optimization With Application to Parameter Tuning for Electronically Controlled Transmissions

RA-L 2020

In mechanical systems, control parameters are often manually tuned by an expert through trial and error, which is labor-intensive and time-consuming. In addition, the difficulty of this problem is that there often exist multiple solutions that provide high returns. As a designed objective function i

Cited by 2SourceScholar
2019

Hierarchical Reinforcement Learning via Advantage-Weighted Information Maximization

ICLR 2019poster

Real-world tasks are often highly structured. Hierarchical reinforcement learning (HRL) has attracted research interest as an approach for leveraging the hierarchical structure of a given task in reinforcement learning (RL). However, identifying the hierarchical policy structure that enhances the pe…

2018

Sample and Feedback Efficient Hierarchical Reinforcement Learning from Human Preferences

ICRA 2018poster

While reinforcement learning has led to promising results in robotics, defining an informative reward function is challenging. Prior work considered including the human in the loop to jointly learn the reward function and the optimal policy. Generating samples from a physical robot and requesting hu…

Cited by 28SourceScholar
2017

A learning-based shared control architecture for interactive task execution

ICRA 2017poster

Shared control is a key technology for various robotic applications in which a robotic system and a human operator are meant to collaborate efficiently. In order to achieve efficient task execution in shared control, it is essential to predict the desired behavior for a given situation or context in…

Cited by 66SourceScholar
2017

Active Incremental Learning of Robot Movement Primitives

CoRL 2017

Robots that can learn over time by interacting with non-technical users must be capable of acquiring new motor skills, incrementally. The problem then is deciding when to teach the robot a new skill or when to rely on the robot generalizing its actions. This decision can be made by the robot if it i

Cited by 0SourcePDFScholar
2017

Guiding Trajectory Optimization by Demonstrated Distributions

RA-L 2017

Trajectory optimization is an essential tool for motion planning under multiple constraints of robotic manipulators. Optimization-based methods can explicitly optimize a trajectory by leveraging prior knowledge of the system and have been used in various applications such as collision avoidance. How

Cited by 60SourceScholar