← Search

Junbo Tan

11 accepted papers

2025

Behavior Cloning Assisted Reinforcement Learning for Cable-Driven Continuum Space Robots in Sparse Reward Environments

RA-L 2025

Deep reinforcement learning (DRL) has emerged as a powerful tool for controlling cable-driven continuum space robots (CDCSRs), offering a solution that bypasses complex system modeling. However, DRL based on dense reward functions (DRLDR) requires meticulous tuning of the reward structure, whereas D

Cited by 1SourceScholar
2025

FOSP: Fine-tuning Offline Safe Policy through World Models

ICLR 2025poster

Offline Safe Reinforcement Learning (RL) seeks to address safety constraints by learning from static datasets and restricting exploration. However, these approaches heavily rely on the dataset and struggle to generalize to unseen scenarios safely. In this paper, we aim to improve safety during the d…

2025

Robust Policy Expansion for Offline-to-Online RL under Diverse Data Corruption

NeurIPS 2025poster

Pretraining a policy on offline data followed by fine-tuning through online interactions, known as Offline-to-Online Reinforcement Learning (O2O RL), has emerged as a promising paradigm for real-world RL deployment. However, both offline datasets and online interactions in practical environments are…

Cited by 0SourcecodeScholar
2024

Hybrid Trajectory Optimization for Autonomous Terrain Traversal of Articulated Tracked Robots

RA-L 2024

Autonomous terrain traversal of articulated tracked robots can reduce operator cognitive load to enhance task efficiency and facilitate extensive deployment. We present a novel hybrid trajectory optimization method aimed at generating efficient, stable, and smooth traversal motions. To achieve this,

Cited by 12SourceScholar
2024

Offline Goal-Conditioned Reinforcement Learning for Safety-Critical Tasks with Recovery Policy

ICRA 2024poster

Offline goal-conditioned reinforcement learning (GCRL) aims at solving goal-reaching tasks with sparse rewards from an offline dataset. While prior work has demonstrated various approaches for agents to learn near-optimal policies, these methods encounter limitations when dealing with diverse constr…

Cited by 6SourcecodeScholar
2023

Visuotactile Sensor Enabled Pneumatic Device Towards Compliant Oropharyngeal Swab Sampling

IROS 2023poster

Manual oropharyngeal (OP) swab sampling is an intensive and risky task. In this article, a novel OP swab sampling device of low cost and high compliance is designed by combining the visuotactile sensor and the pneumatic actuator-based gripper. Here, a concave visuotactile sensor called CoTac is firs…

Cited by 3SourceScholar