← Search

Xueqian Wang

69 accepted papers

2026

CAUSALNAV: A Long-Term Embodied Navigation System for Autonomous Mobile Robots in Dynamic Outdoor Scenarios

RA-L 2026

Autonomous language-guided navigation in large-scale outdoor environments remains a key challenge in mobile robotics, due to difficulties in semantic reasoning, dynamic conditions, and long-term stability. We propose CausalNav, the first scene graph-based semantic navigation framework tailored for d

Cited by 2SourceScholar
2026

CableSense: MuJoCo Simulation-Guided Neural Networks for Force Estimation in Cable-Driven Manipulators

ICRA 2026poster

Cable-driven serial manipulator (CDSM) has advantages of lightweight structure, high flexibility, and inherent safety, making it suitable for operations in constrained spaces. However, interaction with the environment is inevitable. To address this limitation, we propose CableSense, a novel force-se…

Cited by 0Scholar
2026

Diffusion-Enhanced Tree Planning for Autonomous Driving

RA-L 2026

In highly interactive urban driving, decision making is often naturally multi-stage, and decisions at different stages can lead to different reactions from surrounding vehicles. This calls for stage-wise evaluation and selection. Tree-based planning naturally supports multi-stage search and evaluati

Cited by 0SourceScholar
2026

GSON: A Group-Based Social Navigation Framework with Large Multimodal Model

ICRA 2026poster

With the increasing presence of service robots and autonomous vehicles in human environments, navigation systems need to evolve beyond simple destination reach to incorporate social awareness. This paper introduces GSON, a novel group-based social navigation framework that leverages Large Multimodal…

2026

MANSION: Multi-floor lANguage-to-3D Scene generatIOn for loNg-horizon tasks

CVPR 2026

Real-world robotic tasks are long-horizon and often span multiple floors, demanding rich spatial reasoning. However, existing embodied benchmarks are largely confined to single-floor in-house environments, failing to reflect the complexity of real-world tasks. We introduce MANSION, the first languag

Cited by 0SourceScholar
2026

Principled RL for Flow Matching Emerges From the Chunk-level Policy Optimization

ICML 2026poster

Recent Progress in post-training flow matching for text-to-image (T2I) generation with Group Relative Policy Optimization (GRPO) has demonstrated strong potential. However, it is hindered by a critical limitation: inaccurate advantage attribution. In this work, we argue that aggregating consecutive …

Cited by 0SourceScholar
2025

A Novel Split Deep Unfolding Transformer for Pan-Sharpening

ICASSP 2025accepted

Pan-sharpening is a commonly employed strategy to obtain high-resolution multispectral (HRMS) images. Existing deep unfolding networks for pan-sharpening suffer from ineffectively establishing the relationship between panchromatic (PAN) images and generated noisy HRMS (GN-HRMS) images in PAN-guided…

Cited by 0SourceScholar
2025

AirTouch: A Low-Cost Versatile Visuotactile Feedback System for Enhanced Robotic Teleoperation

IROS 2025

Vision-based teleoperation systems are widely used due to their cost-effectiveness and intuitive operation. However, these systems often suffer from challenges such as hand occlusions, environmental variability, and the lack of tactile feedback, limiting their precision and applicability in complex

Cited by 0SourceScholar
2025

Behavior Cloning Assisted Reinforcement Learning for Cable-Driven Continuum Space Robots in Sparse Reward Environments

RA-L 2025

Deep reinforcement learning (DRL) has emerged as a powerful tool for controlling cable-driven continuum space robots (CDCSRs), offering a solution that bypasses complex system modeling. However, DRL based on dense reward functions (DRLDR) requires meticulous tuning of the reward structure, whereas D

Cited by 1SourceScholar
2025

D3-ARM: High-Dynamic, Dexterous and Fully Decoupled Cable-Driven Robotic Arm

ICRA 2025

Cable transmission enables motors of robotic arm to operate lightweight and low-inertia joints remotely in various environments, but it also creates issues with motion coupling and cable routing that can reduce arm's control precision and performance. In this paper, we present a novel motion decoupl

Cited by 5SourceScholar
2025

Entropy-based Activation Function Optimization: A Method on Searching Better Activation Functions

ICLR 2025poster

The success of artificial neural networks (ANNs) hinges greatly on the judicious selection of an activation function, introducing non-linearity into network and enabling them to model sophisticated relationships in data. However, the search of activation functions has largely relied on empirical kno…

Cited by 0SourcePDFScholar
2025

FOSP: Fine-tuning Offline Safe Policy through World Models

ICLR 2025poster

Offline Safe Reinforcement Learning (RL) seeks to address safety constraints by learning from static datasets and restricting exploration. However, these approaches heavily rely on the dataset and struggle to generalize to unseen scenarios safely. In this paper, we aim to improve safety during the d…

2025

GSON: A Group-Based Social Navigation Framework With Large Multimodal Model

RA-L 2025

With the increasing presence of service robots and autonomous vehicles in human environments, navigation systems need to evolve beyond simple destination reach to incorporate social awareness. This paper introduces <sc xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/199

Cited by 10SourceScholar
2025

Generalizing Alignment Paradigm of Text-to-Image Generation with Preferences Through f-Divergence Minimization

AAAI 2025technical

Direct Preference Optimization (DPO) has recently expanded its successful application from aligning large language models (LLMs) to aligning text-to-image models with human preferences, which has generated considerable interest within the community. However, we have observed that these approaches re…

Cited by 4SourcePDFScholar
2025

Identical Human Preference Alignment Paradigm for Text-to-Image Models

ICASSP 2025accepted

Implicit reward mechanism of Direct Preference Optimization (DPO) has facilitated its recent applications beyond large language models (LLMs), notably in aligning text-to-image models with human preferences. While promising results have been achieved with algorithms such as Diffusion-DPO, their reli…

Cited by 0SourceScholar
2025

Learning Generalizable Language-Conditioned Cloth Manipulation from Long Demonstrations

IROS 2025

Multi-step cloth manipulation is a challenging problem for robots due to the high-dimensional state spaces and the dynamics of cloth. Despite recent significant advances in end-to-end imitation learning for multi-step cloth manipulation skills, these methods fail to generalize to unseen tasks. Our i

Cited by 2SourceScholar
2025

Lifelong Safety Alignment for Language Models

NeurIPS 2025poster

LLMs have made impressive progress, but their growing capabilities also expose them to highly flexible jailbreaking attacks designed to bypass safety alignment. While many existing defenses focus on known types of attacks, it is more critical to prepare LLMs for *unseen* attacks that may arise durin…

Cited by 0SourcecodeScholar
2025

Mastering Massive Multi-Task Reinforcement Learning via Mixture-of-Expert Decision Transformer

ICML 2025poster

Despite recent advancements in offline multi-task reinforcement learning (MTRL) have harnessed the powerful capabilities of the Transformer architecture, most approaches focus on a limited number of tasks, with scaling to extremely massive tasks remaining a formidable challenge. In this paper, we f…

2025

MuxHand: A Cost-Effective and Compact Dexterous Robotic Hand Using Time-Division Multiplexing Mechanism

IROS 2025

The number of motors directly influences the dexterity, size, and cost of a robotic hand. In this paper, we present MuxHand, a robotic hand that utilizes a time-division multiplexing motor (TDMM) mechanism. This system enables independent control of 9 cables with just 4 motors, significantly reducin

Cited by 1SourceScholar
2025

Positive Enhanced Preference Alignment for Text-to-Image Models

ICASSP 2025accepted

Direct Preference Optimization (DPO) has recently expanded its successful application beyond aligning large language models (LLMs), further targeting the alignment of text-to-image models with human preferences. However, traditional DPO approach would inadvertently result in a simultaneous reduction…

Cited by 0SourceScholar
2025

Reinforcement Learning Meets Masked Generative Models: Mask-GRPO for Text-to-Image Generation

NeurIPS 2025poster

Reinforcement learning (RL) has garnered increasing attention in text-to-image (T2I) generation. However, most existing RL approaches are tailored to either diffusion models or autoregressive models, overlooking an important alternative: masked generative models. In this work, we propose Mask-GRPO,…

Cited by 0SourceScholar
2025

Robust Policy Expansion for Offline-to-Online RL under Diverse Data Corruption

NeurIPS 2025poster

Pretraining a policy on offline data followed by fine-tuning through online interactions, known as Offline-to-Online Reinforcement Learning (O2O RL), has emerged as a promising paradigm for real-world RL deployment. However, both offline datasets and online interactions in practical environments are…

Cited by 0SourcecodeScholar
2025

Scalable MARL for Cooperative Exploration with Dynamic Robot Populations via Graph-Based Information Aggregation

IROS 2025

This study addresses the challenge of multi-robot cooperative exploration under limited local observations in environments with dynamic robot populations. To achieve efficient area coverage within constrained timeframes, we propose the Multi-Robot Informative Planner (MIP), a novel reinforcement lea

Cited by 0SourceScholar
2025

Tissue-View Map for Robotic Carotid Artery Ultrasound Scanning Using Reinforcement Learning

RA-L 2025

Ultrasound is an important diagnostic modality in medicine, offering real-time imaging, no radiation and low cost. However, ultrasound is currently highly dependent on the operator's experience and technical skills. Robotic autonomous ultrasound scanning (RAUS) is a sequential decision-making proble

Cited by 0SourceScholar
2024

Decentralized Directed Collaboration for Personalized Federated Learning

CVPR 2024poster

Personalized Federated Learning (PFL) is proposed to find the greatest personalized models for each client. To avoid the central failure and communication bottleneck in the server-based FL we concentrate on the Decentralized Personalized Federated Learning (DPFL) that performs distributed model trai…

Cited by 8SourcePDFScholar
2024

Hybrid Trajectory Optimization for Autonomous Terrain Traversal of Articulated Tracked Robots

RA-L 2024

Autonomous terrain traversal of articulated tracked robots can reduce operator cognitive load to enhance task efficiency and facilitate extensive deployment. We present a novel hybrid trajectory optimization method aimed at generating efficient, stable, and smooth traversal motions. To achieve this,

Cited by 12SourceScholar
2024

Learning Language-Conditioned Deformable Object Manipulation with Graph Dynamics

ICRA 2024poster

Multi-task learning of deformable object manipulation is a challenging problem in robot manipulation. Most previous works address this problem in a goal-conditioned way and adapt goal images to specify different tasks, which limits the multi-task learning performance and can not generalize to new ta…

Cited by 14SourceScholar
2024

Offline Goal-Conditioned Reinforcement Learning for Safety-Critical Tasks with Recovery Policy

ICRA 2024poster

Offline goal-conditioned reinforcement learning (GCRL) aims at solving goal-reaching tasks with sparse rewards from an offline dataset. While prior work has demonstrated various approaches for agents to learn near-optimal policies, these methods encounter limitations when dealing with diverse constr…

Cited by 6SourcecodeScholar
2024

Path Generation for Wheeled Robots Autonomous Navigation on Vegetated Terrain

RA-L 2024

Wheeled robot navigation has been widely used in urban environments, but navigation in wild vegetation is still challenging. External sensors (LiDAR, camera etc.) are often used to construct point cloud map of the surrounding environment, however, the supporting rigid ground used for travelling cann

Cited by 29SourceScholar
2024

RH-Map: Online Map Construction Framework of Dynamic Object Removal Based on 3D Region-Wise Hash Map Structure

RA-L 2024

Mobile robots navigating in outdoor environments frequently encounter the issue of undesired traces left by dynamic objects and manifested as obstacles on map, impeding robots from achieving accurate localization and effective navigation. To tackle the problem, a novel map construction framework bas

Cited by 14SourcecodeScholar
2024

Revisiting Plasticity in Visual Reinforcement Learning: Data, Modules and Training Stages

ICLR 2024poster

Plasticity, the ability of a neural network to evolve with new data, is crucial for high-performance and sample-efficient visual reinforcement learning (VRL). Although methods like resetting and regularization can potentially mitigate plasticity loss, the influences of various components within the…

2024

TPTU-v2: Boosting Task Planning and Tool Usage of Large Language Model-based Agents in Real-world Industry Systems

EMNLP 2024industry

Large Language Models (LLMs) have demonstrated proficiency in addressing tasks that necessitate a combination of task planning and the usage of external tools, such as weather and calculator APIs. However, real-world industrial systems present prevalent challenges in task planning and tool usage: nu…

2023

Curriculum-based Co-design of Morphology and Control of Voxel-based Soft Robots

ICLR 2023poster

Co-design of morphology and control of a Voxel-based Soft Robot (VSR) is challenging due to the notorious bi-level optimization. In this paper, we present a Curriculum-based Co-design (CuCo) method for learning to design and control VSRs through an easy-to-difficult process. Specifically, we expand…

Cited by 10SourcePDFScholar
2023

Dynamic Control Barrier Function-based Model Predictive Control to Safety-Critical Obstacle-Avoidance of Mobile Robot

ICRA 2023poster

This paper presents an efficient and safe method to avoid static and dynamic obstacles based on LiDAR. First, point cloud is used to generate a real-time local grid map for obstacle detection. Then, obstacles are clustered by DBSCAN algorithm and enclosed with minimum bounding ellipses (MBEs). In ad…

Cited by 94SourcecodeScholar
2023

Evaluating Model-Free Reinforcement Learning toward Safety-Critical Tasks

AAAI 2023technical

Safety comes first in many real-world applications involving autonomous agents. Despite a large number of reinforcement learning (RL) methods focusing on safety-critical tasks, there is still a lack of high-quality evaluation of those algorithms that adheres to safety constraints at each decision st…

Cited by 31SourcePDFScholar
2023

Foldsformer: Learning Sequential Multi-Step Cloth Manipulation With Space-Time Attention

RA-L 2023

Sequential multi-step cloth manipulation is a challenging problem in robotic manipulation, requiring a robot to perceive the cloth state and plan a sequence of chained actions leading to the desired state. Most previous works address this problem in a goal-conditioned way, and goal observation must

Cited by 33SourcecodeScholar
2023

Improving the Model Consistency of Decentralized Federated Learning

ICML 2023poster

To mitigate the privacy leakages and communication burdens of Federated Learning (FL), decentralized FL (DFL) discards the central server and each client only communicates with its neighbors in a decentralized communication network. However, existing DFL suffers from high inconsistency among local c…

Cited by 68SourcePDFScholar
2023

Learning Better with Less: Effective Augmentation for Sample-Efficient Visual Reinforcement Learning

NeurIPS 2023poster

Data augmentation (DA) is a crucial technique for enhancing the sample efficiency of visual reinforcement learning (RL) algorithms. Notably, employing simple observation transformations alone can yield outstanding performance without extra auxiliary representation tasks or pre-trained encoders. Howe…

2023

Learning Graph Dynamics With External Contact for Deformable Linear Objects Shape Control

RA-L 2023

This letter focuses on the shape control manipulation of deformable linear objects (DLO) with a dual-arm robotic system. One significant challenge of DLO shape control is the underactuated control system, which means that finite robotic manipulators can not fully control DLO's shape due to the lack

Cited by 19SourceScholar
2023

Make Landscape Flatter in Differentially Private Federated Learning

CVPR 2023poster

To defend the inference attacks and mitigate the sensitive information leakages in Federated Learning (FL), client-level Differentially Private FL (DPFL) is the de-facto standard for privacy protection by clipping local updates and adding random noise. However, existing DPFL methods tend to make a s…

2023

PreCo: Enhancing Generalization in Co-Design of Modular Soft Robots via Brain-Body Pre-Training

CoRL 2023oral

Brain-body co-design, which involves the collaborative design of control strategies and morphologies, has emerged as a promising approach to enhance a robot's adaptability to its environment. However, the conventional co-design process often starts from scratch, lacking the utilization of prior know…

Cited by 9SourceScholar
2023

Quadruped Guidance Robot for the Visually Impaired: A Comfort-Based Approach

ICRA 2023poster

Guidance robots that can guide people and avoid various obstacles, could potentially be owned by more visually impaired people at a fairly low cost. Most of the previous guidance robots for the visually impaired ignored the human response behavior and comfort, treating the human as an appendage drag…

Cited by 45SourceScholar
2023

TEFISTA-NET: GTD Parameter Estimation of Low-Frequency Ultra- Wideband Radar via Model-Based Deep Learning

ICASSP 2023accepted

The geometrical theory of diffraction (GTD) has been widely investigated to describe the target scattering behaviors with the low-frequency ultra-wideband (LFW) radar. In this paper, we propose a new model-based deep learning method for GTD parameter estimation. The proposed method is designed by un…

Cited by 0SourceScholar
2023

Transparent Shape from a Single View Polarization Image

ICCV 2023poster

This paper presents a learning-based method for transparent surface estimation from a single view polarization image. Existing shape from polarization(SfP) methods have the difficulty in estimating transparent shape since the inherent transmission interference heavily reduces the reliability of phys…

Cited by 11PDFcodeScholar
2023

USEEK: Unsupervised SE(3)-Equivariant 3D Keypoints for Generalizable Manipulation

ICRA 2023poster

Can a robot manipulate intra-category unseen objects in arbitrary poses with the help of a mere demonstration of grasping pose on a single object instance? In this paper, we try to address this intriguing challenge by using USEEK, an unsupervised SE(3)-equivariant keypoints method that enjoys alignm…

Cited by 29SourceScholar
2023

Visuotactile Sensor Enabled Pneumatic Device Towards Compliant Oropharyngeal Swab Sampling

IROS 2023poster

Manual oropharyngeal (OP) swab sampling is an intensive and risky task. In this article, a novel OP swab sampling device of low cost and high compliance is designed by combining the visuotactile sensor and the pneumatic actuator-based gripper. Here, a concave visuotactile sensor called CoTac is firs…

Cited by 3SourceScholar
2023

Volumetric 3D Reconstruction with Window-Wise Global Feature Aggregation

ICASSP 2023accepted

Volumetric 3D reconstruction methods have shown great performance in reconstructing indoor scenarios from monocular videos. However, as such approaches utilize discrete feature voxels to encode the observed scenes, the global feature interaction within and across different voxels is ignored, leading…

Cited by 0SourceScholar
2022

Deep Reinforcement Learning Based on Local GNN for Goal-Conditioned Deformable Object Rearranging

IROS 2022poster

Object rearranging is one of the most common deformable manipulation tasks, where the robot needs to rearrange a deformable object into a goal configuration. Previous studies focus on designing an expert system for each specific task by model-based or data-driven approaches and the application scena…

Cited by 17SourceScholar
2022

Don’t Touch What Matters: Task-Aware Lipschitz Data Augmentation for Visual Reinforcement Learning

IJCAI 2022poster

One of the key challenges in visual Reinforcement Learning (RL) is to learn policies that can generalize to unseen environments. Recently, data augmentation techniques aiming at enhancing data diversity have demonstrated proven performance in improving the generalization ability of learned policies.…

2022

Orientation to Pose: Continuum Robots Shape Reconstruction Based on the Multi-Attitude Solving Approach

ICRA 2022poster

Continuum robots are typically slender and flexible with infinite freedoms in theory, which poses a challenge for their control and application. The shape reconstruction of continuum robots is vital to realize closed-loop control. This paper proposes a novel general real-time shape reconstruction fr…

Cited by 7SourceScholar
2022

PUTN: A Plane-fitting based Uneven Terrain Navigation Framework

IROS 2022poster

Autonomous navigation of ground robots has been widely used in indoor structured 2D environments, but there are still many challenges in outdoor 3D unstructured environments, especially in rough, uneven terrains. This paper proposed a plane-fitting based uneven terrain navigation framework (PUTN) to…

Cited by 59SourcecodeScholar
2022

Penalized Proximal Policy Optimization for Safe Reinforcement Learning

IJCAI 2022poster

Safe reinforcement learning aims to learn the optimal policy while satisfying safety constraints, which is essential in real-world applications. However, current algorithms still struggle for efficient policy updates with hard constraint satisfaction. In this paper, we propose Penalized Proximal Pol…

2022

Pre-Trained Image Encoder for Generalizable Visual Reinforcement Learning

NeurIPS 2022accept

Learning generalizable policies that can adapt to unseen environments remains challenging in visual Reinforcement Learning (RL). Existing approaches try to acquire a robust representation via diversifying the appearances of in-domain observations for better generalization. Limited by the specific ob…

Cited by 85SourcePDFScholar
2022

Safety Correction from Baseline: Towards the Risk-aware Policy in Robotics via Dual-agent Reinforcement Learning

IROS 2022poster

Learning a risk-aware policy is essential but rather challenging in unstructured robotic tasks. Safe reinforcement learning methods open up new possibilities to tackle this problem. However, the conservative policy updates make it intractable to achieve sufficient exploration and desirable performan…

Cited by 4SourceScholar
2022

TaTa: A Universal Jamming Gripper with High-Quality Tactile Perception and Its Application to Underwater Manipulation

ICRA 2022poster

Large-area and high-precision tactile sensing information can not only improve the stability of robot grasping but also compensate for the lack of visual information in specific environments such as turbid underwater, dimness, and smoke. In this paper, we devise a universal jamming gripper with high…

Cited by 33SourceScholar
2022

Vertebraic Soft Robotic Joint Design With Twisting and Antagonism

RA-L 2022

The soft robotic manipulators attract extensive interest of researchers due to its conformity to the unstructured environment, safe-interaction with human and fragile objects. The movement of the soft manipulator often include elongation, contraction, 2-DOF rotations due to the parallelly arranged f

Cited by 16SourceScholar
2021

An Overall Configuration Planning Method of Continuum Hyper-Redundant Manipulators Based on Improved Artificial Potential Field Method

RA-L 2021

Continuum hyper-redundant manipulators (CHRMs) have been widely applied in aerospace, medical or other fields to complete tasks in narrow and multi-obstacles environments with its unique structural advantages. Due to the redundancy, the inverse kinematics of CHRMs is rather complex and the trajector

Cited by 46SourceScholar
2021

Soft-CCD Algorithm for Inverse Kinematics of Soft Continuum Manipulators

IROS 2021poster

To date, soft robots have been increasingly designed and analyzed, especially, Soft Continuum Manipulators (SCMs). Due to dexterous deformability, their Inverse Kinematics (IK) is still difficult to solve. Cyclic Coordinate Descent (CCD) algorithm is one of the classical optimization algorithms to s…

Cited by 6SourceScholar
2020

Modeling and Experimental Verification of a Cable-Constrained Synchronous Rotating Mechanism Considering Friction Effect

RA-L 2020

Cable-Constrained Synchronous Rotating Mechanism (CCSRM) has an important application prospect in the field of cable-driven robots, which can greatly reduce the number of driving motors while ensuring the light and slender body. However, there are obvious cable friction effect and elastic deformatio

Cited by 16SourceScholar
2020

Multi-task Control for a Quadruped Robot with Changeable Leg Configuration

IROS 2020poster

This paper proposes a multi-task control strategy for a quadruped robot named THU-QUAD II. The mechanical design of the robot ensures a wide range of motion for all joints, which allows it to stand and walk like a mammal as well as sprawl to the ground and crawl like a reptile. Five basic leg config…

Cited by 8SourceScholar
2019

A 3D Static Modeling Method and Experimental Verification of Continuum Robots Based on Pseudo-Rigid Body Theory

IROS 2019poster

Continuum robots composed of elastic backbones have a broad application prospect in the narrow and restricted environment because they overcome the disadvantages of traditional articulated robots, such as being bulky and inflexible. Statics plays an important role in the planning and control of the…

Cited by 35SourceScholar