← Search

Yan Wu

51 accepted papers

2026

Effective Robotic Cloth Grasping Through Suppressing False Discoveries

AAAI 2026technical

Enabling robots to grasp disorganized cloth for efficient storage is valuable in robot-assisted room organization. Diverse deformations of cloth and the stacking of multiple items limit grasping-pose estimation that relies on annotations. This necessitates segmenting each cloth item in an unsupervis

Cited by 0SourcePDFScholar
2026

Guiding Robotic Cloth Grasping in Darkness: Infrared Semantic Segmentation and Grasping Position Selection

RA-L 2026

Robotic cloth grasping is a key component in many robotic cloth manipulation scenarios, such as automated wardrobe management, clothing laundering, and assisted dressing. Due to the deformability and large surface of cloth, which distinguishes it from conventional rigid targets, most current studies

Cited by 2SourceScholar
2026

Harmonising Safety Paradigms: Energy-Aware Control of Active Response and Passive Compliance for Safety-Critical Robotic Tasks

RA-L 2026

Ensuring safety in robotic manipulation is increasingly critical as robots become integrated into human-shared environments for complex physical interaction tasks. This paper presents an energy-aware control framework that combines active responses with passive compliance for safety-critical robotic

Cited by 0SourceScholar
2026

Harmonising Safety Paradigms: Energy-Aware Control of Active Response and Passive Compliance for Safety-Critical Robotic Tasks

ICRA 2026poster

Ensuring safety in robotic manipulation is increasingly critical as robots become integrated into human-shared environments for complex physical interaction tasks. This paper presents an energy-aware control framework that combines active responses with passive compliance for safety-critical robotic…

Cited by 0SourceScholar
2026

MachaGrasp: Morphology-Aware Cross-Embodiment Dexterous Hand Articulation Generation for Grasping

ICRA 2026poster

Dexterous grasping with multi-fingered hands remains challenging due to high-dimensional articulations and the cost of optimization-based pipelines. Existing end-to-end methods require training on large-scale datasets for specific hands, limiting their ability to generalize across different embodime…

2026

Semantic Contact Fields for Category-Level Generalizable Tool Manipulation

RSS 2026poster

Generalizing tool manipulation requires both semantic planning and precise physical control. Modern generalist robot policies, such as Vision-Language-Action (VLA) models, often lack the high-fidelity physical grounding required for contact-rich tool manipulation. Conversely, existing contact-aware …

Cited by 0SourceScholar
2025

CPA-Enhancer: Chain-of-Thought Prompted Adaptive Enhancer for Downstream Vision Tasks Under Unknown Degradations

ICASSP 2025accepted

Extracting valuable visual cues for downstream vision tasks poses a particular challenge under unknown degradations. A straightforward solution is to preprocess images using image restoration methods, but their high computational complexity renders them unsuitable for real-time tasks. Recent efforts…

Cited by 0SourceScholar
2025

DiMSOD: A Diffusion-Based Framework for Multi-Modal Salient Object Detection

AAAI 2025technical

Multi-modal salient object detection (SOD) through the integration of additional data such as depth or thermal information has become a significant task in computer vision during recent years. Traditionally, the challenges of identifying salient objects in RGB, RGB-D (Depth), and RGB-T (Thermal) ima…

Cited by 0SourcePDFScholar
2025

Generalizable Category-Level Topological Structure Learning for Clothing Recognition in Robotic Grasping

IROS 2025

Recognizing various types of clothing is crucial for robotic clothing manipulation tasks, such as garment organization and robot-assisted dressing. Unlike rigid object recognition, clothing recognition remains a challenging task due to the diverse forms introduced by flexible deformations. However,

Cited by 1SourceScholar
2025

LEMoN: Label Error Detection using Multimodal Neighbors

ICML 2025poster

Large repositories of image-caption pairs are essential for the development of vision-language models. However, these datasets are often extracted from noisy data scraped from the web, and contain many mislabeled instances. In order to improve the reliability of downstream models, it is important to…

Cited by 0SourcePDFScholar
2025

Learning-Based Predictive Impedance Control Towards Safe Predefined-Time Physical Robotic Interaction

IROS 2025

Impedance control can be achieved within a model predictive control (MPC) framework for optimization and constraint compliance. However, user-defined or optimization-derived impedance models can be too conservative to achieve a timely convergence, or too aggressive to ensure safety. To address this,

Cited by 0SourceScholar
2025

PerturBench: Benchmarking Machine Learning Models for Cellular Perturbation Analysis

NeurIPS 2025poster

We introduce a comprehensive framework for modeling single cell transcriptomic responses to perturbations, aimed at standardizing benchmarking in this rapidly evolving field. Our approach includes a modular and user-friendly model development and evaluation platform, a collection of diverse perturba…

Cited by 0SourcecodeScholar
2025

Seg-diffusion: Text-to-Image Diffusion Model for Open-Vocabulary Semantic Segmentation

ICASSP 2025accepted

Open-vocabulary semantic segmentation (OVSS) is a challenging computer vision task that labels each pixel within an image based on text descriptions. Recent advancements in OVSS are largely attributed to the increased model capacity. However, these models often struggle with unfamiliar images or uns…

Cited by 0SourceScholar
2025

Sequen-Sync Contact Force/Torque Control Using Nested Fast Terminal Sliding Mode Control Approach

IROS 2025

As one of the most fundamental control modes in robotics, force/torque (F/T) control plays an essential role in a wide range of applications. However, classical F/T control fails to offer effective means to regulate the convergence sequence of the controlled states, which is beneficial in many real-

Cited by 0SourceScholar
2025

UniPhys: Unified Planner and Controller with Diffusion for Flexible Physics-Based Character Control

ICCV 2025poster

Generating natural and physically plausible character motion remains challenging, particularly for long-horizon control with diverse guidance signals. While prior work combines high-level diffusion-based motion planners with low-level physics controllers, these systems suffer from domain gaps that d…

Cited by 0SourcePDFScholar
2024

Protein Discovery with Discrete Walk-Jump Sampling

ICLR 2024oral

We resolve difficulties in training and sampling from a discrete generative model by learning a smoothed energy function, sampling from the smoothed data manifold with Langevin Markov chain Monte Carlo (MCMC), and projecting back to the true data manifold with one-step denoising. Our $\textit{Discre…

2024

Unknown Object Retrieval in Confined Space through Reinforcement Learning with Tactile Exploration

ICRA 2024poster

The potential of tactile sensing for dexterous robotic manipulation has been demonstrated by its ability to enable nuanced real-world interactions. In this study, the retrieval of unknown objects from confined spaces, which is unsuitable for conventional visual perception and gripper-based manipulat…

Cited by 0SourceScholar
2024

Visual-Policy Learning Through Multi-Camera View to Single-Camera View Knowledge Distillation for Robot Manipulation Tasks

RA-L 2024

The use of multi-camera views simultaneously has been shown to improve the generalization capabilities and performance of visual policies. However, using multiple cameras in real-world scenarios can be challenging. In this study, we present a novel approach to enhance the generalization performance

Cited by 11SourceScholar
2023

AbDiffuser: full-atom generation of in-vitro functioning antibodies

NeurIPS 2023spotlight

We introduce AbDiffuser, an equivariant and physics-informed diffusion model for the joint generation of antibody 3D structures and sequences. AbDiffuser is built on top of a new representation of protein structure, relies on a novel architecture for aligned proteins, and utilizes strong diffusion p…

Cited by 52SourcePDFScholar
2023

Interaction Control for Tool Manipulation on Deformable Objects Using Tactile Feedback

RA-L 2023

The human sense of touch enables us to perform delicate tasks on deformable objects and/or in a vision-denied environment. To achieve similar desirable interactions for robots, such as administering a swab test, tactile information sensed beyond the tool-in-hand is crucial for contact state estimati

Cited by 9SourceScholar
2023

Learning Deep Sensorimotor Policies for Vision-Based Autonomous Drone Racing

IROS 2023poster

The development of effective vision-based algorithms has been a significant challenge in achieving autonomous drones, which promise to offer immense potential for many real-world applications. This paper investigates learning deep sensorimotor policies for vision-based drone racing, which is a parti…

Cited by 21SourceScholar
2023

Visuo-Tactile Feedback-Based Robot Manipulation for Object Packing

RA-L 2023

Robots are increasingly expected to manipulate objects, of which properties have high perceptual uncertainty from any single sensory modality. This directly impacts successful object manipulation. Object packing is one of the challenging tasks in robot manipulation. In this work, a new visuo-tactile

Cited by 23SourceScholar
2022

CRAFT: Cross-Attentional Flow Transformer for Robust Optical Flow

CVPR 2022poster

Optical flow estimation aims to find the 2D motion field by identifying corresponding pixels between two images. Despite the tremendous progress of deep learning-based optical flow methods, it remains a challenge to accurately estimate large displacements with motion blur. This is mainly because the…

Cited by 134PDFcodeScholar
2022

SAGA: Stochastic Whole-Body Grasping with Contact

ECCV 2022poster

"The synthesis of human grasping has numerous applications including AR/VR, video games and robotics. While methods have been proposed to generate realistic hand-object interaction for object grasping and manipulation, these typically only consider interacting hand alone. Our goal is to synthesize w…

2022

Towards Learning Universal Audio Representations

ICASSP 2022accepted

The ability to learn universal audio representations that can solve diverse speech, music, and environment tasks can spur many applications that require general sound content understanding. In this work, we introduce a holistic audio representation evaluation suite (HARES) spanning 12 downstream tas…

Cited by 0SourceScholar
2021

GPU-Efficient Dense Convolutional Network for Real-time Semantic Segmentation

ICRA 2021poster

Real-time semantic segmentation is a challenging task as both accuracy and inference speed need to be considered simultaneously. In real-world applications, it is usually achieved by deploying a deep neural network in modern GPU device. However, most of the work focused on real-time semantic segment…

Cited by 1SourceScholar
2021

Kanerva++: Extending the Kanerva Machine With Differentiable, Locally Block Allocated Latent Memory

ICLR 2021poster

Episodic and semantic memory are critical components of the human memory model. The theory of complementary learning systems (McClelland et al., 1995) suggests that the compressed representation produced by a serial event (episodic memory) is later restructured to build a more generalized form of re…

Cited by 4SourcePDFScholar
2021

Neural Architecture Search of SPD Manifold Networks

IJCAI 2021poster

In this paper, we propose a new neural architecture search (NAS) problem of Symmetric Positive Definite (SPD) manifold networks, aiming to automate the design of SPD neural architectures. To address this problem, we first introduce a geometrically rich and diverse SPD neural architecture search spac…

2021

On Explainability and Sensor-Adaptability of a Robot Tactile Texture Representation Using a Two-Stage Recurrent Networks

IROS 2021poster

The ability to simultaneously distinguish objects, materials, and their associated physical properties is one fundamental function of the sense of touch. Recent advances in the development of tactile sensors and machine learning techniques allow more accurate and complex modelling of robotic tactile…

Cited by 7SourceScholar
2021

Towards Efficient Multiview Object Detection with Adaptive Action Prediction

ICRA 2021poster

Active vision is a desirable perceptual feature for robots. Existing approaches usually make strong assumptions about the task and environment, thus are less robust and efficient. This study proposes an adaptive view planning approach to boost the efficiency and robustness of active object detection…

Cited by 9SourceScholar
2020

Fast Texture Classification Using Tactile Neural Coding and Spiking Neural Network

IROS 2020poster

Touch is arguably the most important sensing modality in physical interactions. However, tactile sensing has been largely under-explored in robotics applications owing to the complexity in making perceptual inferences until the recent advancements in machine learning or deep learning in particular.…

Cited by 35SourceScholar
2020

Multi-UAV Coverage Path Planning for the Inspection of Large and Complex Structures

IROS 2020poster

We present a multi-UAV Coverage Path Planning (CPP) framework for the inspection of large-scale, complex 3D structures. In the proposed sampling-based coverage path planning method, we formulate the multi-UAV inspection applications as a multi-agent coverage path planning problem. By combining two N…

Cited by 84SourceScholar
2020

Robust Force Tracking Impedance Control of an Ultrasonic Motor-actuated End-effector in a Soft Environment

IROS 2020poster

Robotic systems are increasingly required not only to generate precise motions to complete their tasks but also to handle the interactions with the environment or human. Significantly, soft interaction brings great challenges on the force control due to the nonlinear, viscoelastic and inhomogeneous…

Cited by 10SourceScholar
2020

Supervised Autoencoder Joint Learning on Heterogeneous Tactile Sensory Data: Improving Material Classification Performance

IROS 2020poster

The sense of touch is an essential sensing modality for a robot to interact with the environment as it provides rich and multimodal sensory information upon contact. It enriches the perceptual understanding of the environment and closes the loop for action generation. One fundamental area of percept…

Cited by 19SourceScholar
2020

Training Generative Adversarial Networks by Solving Ordinary Differential Equations

NeurIPS 2020spotlight

The instability of Generative Adversarial Network (GAN) training has frequently been attributed to gradient descent. Consequently, recent methods have aimed to tailor the models and training procedures to stabilise the discrete updates. In contrast, we study the continuous-time dynamics induced by G…

2019

DFNet: Semantic Segmentation on Panoramic Images with Dynamic Loss Weights and Residual Fusion Block

ICRA 2019poster

For the domain of self-driving and automatic parking, perception is a basic and critical technique, moreover, the detection of lane markings and parking slots is an important part of visual perception. Compared with front sight images, panoramic images(PI) can capture more comprehensive pavement inf…

Cited by 27SourceScholar
2019

Shaping Belief States with Generative Environment Models for RL

NeurIPS 2019poster

When agents interact with a complex environment, they must form and maintain beliefs about the relevant aspects of that environment. We propose a way to efficiently train expressive generative models in complex environments. We show that a predictive algorithm with an expressive generative model can…

Cited by 127SourcePDFScholar
2019

Towards Effective Tactile Identification of Textures using a Hybrid Touch Approach

ICRA 2019poster

The sense of touch is arguably the first human sense to develop. Empowering robots with the sense of touch may augment their understanding of interacted objects and the environment beyond standard sensory modalities (e.g., vision). This paper investigates the effect of hybridizing touch and sliding…

Cited by 52SourceScholar
2018

Learning Attractor Dynamics for Generative Memory

NeurIPS 2018poster

A central challenge faced by memory systems is the robust retrieval of a stored pattern in the presence of interference due to other stored patterns and noise. A theoretically well-founded solution to robust retrieval is given by attractor dynamics, which iteratively cleans up patterns during recall…

2018

Multi-Modal Robot Apprenticeship: Imitation Learning Using Linearly Decayed DMP+ in a Human-Robot Dialogue System

IROS 2018poster

Robot learning by demonstration gives robots the ability to learn tasks which they have not been programmed to do before. The paradigm allows robots to work in a greater range of real-world applications in our daily life. However, this paradigm has traditionally been applied to learn tasks from a si…

Cited by 25SourceScholar
2016

Dynamic Movement Primitives Plus: For enhanced reproduction quality and efficient trajectory modification using truncated kernels and Local Biases

IROS 2016poster

Dynamic Movement Primitives (DMPs) are a generic approach for trajectory modeling in an attractor land-scape based on differential dynamical systems. DMPs guarantee stability and convergence properties of learned trajectories, and scale well to high dimensional data. In this paper, we propose DMP+,…

Cited by 57SourceScholar
2015

Adaptive optimal control for coordination in physical human-robot interaction

IROS 2015poster

In this paper, we propose an adaptive optimal control for a robot to collaborate with a human. Game theory and policy iteration are employed to analyze the interactive behaviors of the human and the robot in physical interactions. The human's control objective is estimated and it is used to adapt th…

Cited by 26SourceScholar
2015

Intention detection in upper limb kinematics rehabilitation using a GP-based control strategy

IROS 2015poster

In robot-assisted upper limb rehabilitation, detecting the intentions of hemiplegic patients is essential towards assisting the patients to actively exercise instead of driving passive motions. Many interactive channels, such as voice, EMG and EEG, have been studied to estimate the motion intentions…

Cited by 8SourceScholar