← Search

Yixing Gao

17 accepted papers

2026

Effective Robotic Cloth Grasping Through Suppressing False Discoveries

AAAI 2026technical

Enabling robots to grasp disorganized cloth for efficient storage is valuable in robot-assisted room organization. Diverse deformations of cloth and the stacking of multiple items limit grasping-pose estimation that relies on annotations. This necessitates segmenting each cloth item in an unsupervis

Cited by 0SourcePDFScholar
2026

FlatLab: A Unified Methodology Framework and Simulation-Based Benchmark for Robotic Manipulation of Flat Objects

ICML 2026poster

Robotic manipulation of flat objects is challenging due to the ungraspable configurations and strong variations in object geometry and material. Existing methods rely on heuristic pre-manipulation and are often evaluated in closed settings with limited generalization. We propose a unified framework …

Cited by 0SourcecodeScholar
2026

GraspALL: Adaptive Structural Compensation from Illumination Variation for Robotic Garment Grasping in Any Low-Light Conditions

CVPR 2026

Achieving accurate garment grasping under dynamically changing illumination is crucial for all-day operation of service robots. However, the reduced illumination in low-light scenes severely degrades garment structural features, leading to a significant drop in grasping robustness. Existing methods

Cited by 0SourcecodeScholar
2026

Guiding Robotic Cloth Grasping in Darkness: Infrared Semantic Segmentation and Grasping Position Selection

RA-L 2026

Robotic cloth grasping is a key component in many robotic cloth manipulation scenarios, such as automated wardrobe management, clothing laundering, and assisted dressing. Due to the deformability and large surface of cloth, which distinguishes it from conventional rigid targets, most current studies

Cited by 2SourceScholar
2025

AMDANet: Attention-Driven Multi-Perspective Discrepancy Alignment for RGB-Infrared Image Fusion and Segmentation

ICCV 2025poster

The challenge of multimodal semantic segmentation lies in establishing semantically consistent and segmentable multimodal fusion features under conditions of significant visual feature discrepancies. Existing methods commonly construct cross-modal self-attention fusion frameworks or introduce additi…

2025

Collaborative Learning for 3D Hand-Object Reconstruction and Compositional Action Recognition from Egocentric RGB Videos Using Superquadrics

AAAI 2025technical

With the availability of egocentric 3D hand-object interaction datasets, there is increasing interest in developing unified models for hand-object pose estimation and action recognition. However, existing methods still struggle to recognise seen actions on unseen objects due to the limitations in re…

Cited by 0SourcePDFScholar
2025

DarkSeg: Infrared-Driven Semantic Segmentation for Garment Grasping Detection in Low-Light Conditions

IROS 2025

Garment grasping in low-light environments is a critical challenge for domestic intelligent robots, yet existing research has not sufficiently addressed this issue. In low-light conditions, the scarcity of visual features due to insufficient illumination causes different categories of garments to ex

Cited by 1SourcecodeScholar
2025

Generalizable Category-Level Topological Structure Learning for Clothing Recognition in Robotic Grasping

IROS 2025

Recognizing various types of clothing is crucial for robotic clothing manipulation tasks, such as garment organization and robot-assisted dressing. Unlike rigid object recognition, clothing recognition remains a challenging task due to the diverse forms introduced by flexible deformations. However,

Cited by 1SourceScholar
2025

High-Resolution Spatiotemporal Modeling with Global-Local State Space Models for Video-Based Human Pose Estimation

ICCV 2025poster

Modeling high-resolution spatiotemporal representations, including both global dynamic contexts (e.g., holistic human motion tendencies) and local motion details (e.g., high-frequency changes of keypoints), is essential for video-based human pose estimation (VHPE). Current state-of-the-art methods t…

Cited by 0SourcePDFScholar
2024

JointLoc: A Real-time Visual Localization Framework for Planetary UAVs Based on Joint Relative and Absolute Pose Estimation

IROS 2024poster

Unmanned aerial vehicles (UAVs) visual localization in planetary aims to estimate the absolute pose of the UAV in the world coordinate system through satellite maps and images captured by on-board cameras. However, since planetary scenes often lack significant landmarks and there are modal differenc…

Cited by 6SourcecodeScholar
2024

Revealing the Two Sides of Data Augmentation: An Asymmetric Distillation-based Win-Win Solution for Open-Set Recognition

IJCAI 2024poster

In this paper, we reveal the two sides of data augmentation: enhancements in closed-set recognition correlate with a significant decrease in open-set recognition. Through empirical investigation, we find that multi-sample-based augmentations would contribute to reducing feature discrimination, there…

Cited by 1SourcePDFScholar
2023

Clothes Grasping and Unfolding Based on RGB-D Semantic Segmentation

ICRA 2023poster

Clothes grasping and unfolding is a core step in robotic-assisted dressing. Most existing works leverage depth images of clothes to train a deep learning-based model to recognize suitable grasping points. These methods often utilize physics engines to synthesize depth images to reduce the cost of re…

Cited by 5SourceScholar
2023

DiffPose: SpatioTemporal Diffusion Model for Video-Based Human Pose Estimation

ICCV 2023poster

Denoising diffusion probabilistic models that were initially proposed for realistic image generation have recently shown success in various perception tasks (e.g., object detection and image segmentation) and are increasingly gaining attention in computer vision. However, extending such models to mu…

Cited by 43PDFScholar
2023

Mutual Information-Based Temporal Difference Learning for Human Pose Estimation in Video

CVPR 2023poster

Temporal modeling is crucial for multi-frame human pose estimation. Most existing methods directly employ optical flow or deformable convolution to predict full-spectrum motion fields, which might incur numerous irrelevant cues, such as a nearby person or background. Without further efforts to excav…

Cited by 24SourcePDFScholar
2022

Temporal Feature Alignment and Mutual Information Maximization for Video-Based Human Pose Estimation

CVPR 2022oral

Multi-frame human pose estimation has long been a compelling and fundamental problem in computer vision. This task is challenging due to fast motion and pose occlusion that frequently occur in videos. State-of-the-art methods strive to incorporate additional visual evidences from neighboring frames…

Cited by 82PDFcodeScholar
2016

Iterative path optimisation for personalised dressing assistance using vision and force information

IROS 2016poster

We propose an online iterative path optimisation method to enable a Baxter humanoid robot to assist human users to dress. The robot searches for the optimal personalised dressing path using vision and force sensor information: vision information is used to recognise the human pose and model the move…

Cited by 82SourceScholar