← Search

Yang Cong

23 accepted papers

2026

GAPG: Geometry Aware Push-Grasping Synergy for Goal-Oriented Manipulation in Clutter

ICRA 2026poster

Grasping target objects is a fundamental skill for robotic manipulation, but in cluttered environments with stacked or occluded objects, a single-step grasp is often insufficient. To address this, previous work has introduced pushing as an auxiliary action to create graspable space. However, these m…

2026

PixelVLA: Advancing Pixel-level Understanding in Vision-Language-Action Model

ICLR 2026poster

Vision-Language-Action models (VLAs) are emerging as powerful tools for learning generalizable visuomotor control policies. However, current VLAs are mostly trained on large-scale image–text–action data and remain limited in two key ways: (i) they struggle with pixel-level scene understanding, and (…

Cited by 0SourceScholar
2025

Learning Generalizable 3D Manipulation With 10 Demonstrations

IROS 2025

Learning robust and generalizable manipulation skills from few demonstrations remains a key challenge in robotics, with broad applications in industrial automation and service robotics. Although recent imitation learning methods have achieved impressive results, they often require a large amount of

Cited by 2SourcecodeScholar
2024

Cs2K: Class-specific and Class-shared Knowledge Guidance for Incremental Semantic Segmentation

ECCV 2024poster

"Incremental semantic segmentation endeavors to segment newly encountered classes while maintaining knowledge of old classes. However, existing methods either 1) lack guidance from class-specific knowledge (i.e., old class prototypes), leading to a bias towards new classes, or 2) constrain class-sha…

Cited by 1SourcePDFScholar
2023

Augmented Box Replay: Overcoming Foreground Shift for Incremental Object Detection

ICCV 2023poster

In incremental learning, replaying stored samples from previous tasks together with current task samples is one of the most efficient approaches to address catastrophic forgetting. However, unlike incremental classification, image replay has not been successfully applied to incremental object detect…

Cited by 32PDFcodeScholar
2023

Autonomous Manipulation Learning for Similar Deformable Objects via Only One Demonstration

CVPR 2023poster

In comparison with most methods focusing on 3D rigid object recognition and manipulation, deformable objects are more common in our real life but attract less attention. Generally, most existing methods for deformable object manipulation suffer two issues, 1) Massive demonstration: repeating thousan…

Cited by 6SourcePDFScholar
2023

Federated Incremental Semantic Segmentation

CVPR 2023poster

Federated learning-based semantic segmentation (FSS) has drawn widespread attention via decentralized training on local clients. However, most FSS models assume categories are fxed in advance, thus heavily undergoing forgetting on old categories in practical applications where local clients receive…

2022

Class-Incremental Gesture Recognition Learning with Out-of-Distribution Detection

IROS 2022poster

Gesture recognition is a popular human-computer interaction technology, which has been widely applied in many fields (e.g., autonomous driving, medical care, VR and AR). However, 1) most existing gesture recognition methods focus on the fixed recognition scenarios with several gestures, which could…

Cited by 13SourceScholar
2022

The Devil Is in the Pose: Ambiguity-Free 3D Rotation-Invariant Learning via Pose-Aware Convolution

CVPR 2022poster

Recent progress in introducing rotation invariance (RI) to 3D deep learning methods is mainly made by designing RI features to replace 3D coordinates as input. The key to this strategy lies in how to restore the global information that is lost by the input RI features. Most state-of-the-arts achieve…

Cited by 28PDFcodeScholar
2021

Generative Partial Visual-Tactile Fused Object Clustering

AAAI 2021technical

Visual-tactile fused sensing for object clustering has achieved significant progresses recently, since the involvement of tactile modality can effectively improve clustering performance. However, the missing data (i.e., partial data) issues always happen due to occlusion and noises during the data c…

Cited by 17SourcePDFScholar
2021

Unsupervised Dense Deformation Embedding Network for Template-Free Shape Correspondence

ICCV 2021poster

Shape correspondence from 3D deformation learning has attracted appealing academy interests recently. Nevertheless, current deep learning based methods require the supervision of dense annotations to learn per-point translations, which severely over-parameterize the deformation process. Moreover, th…

Cited by 7PDFScholar
2020

CSCL: Critical Semantic-Consistent Learning for Unsupervised Domain Adaptation

ECCV 2020poster

Unsupervised domain adaptation without consuming annotation process for unlabeled target data attracts appealing interests in semantic segmentation. However, 1) existing methods neglect that not all semantic representations across domains are transferable, which cripples domain-wise transfer with un…

Cited by 56SourcePDFScholar
2020

What Can Be Transferred: Unsupervised Domain Adaptation for Endoscopic Lesions Segmentation

CVPR 2020poster

Unsupervised domain adaptation has attracted growing research attention on semantic segmentation. However, 1) most existing models cannot be directly applied into lesions transfer of medical images, due to the diverse appearances of same lesion among different datasets; 2) equal attention has been p…

Cited by 173PDFScholar
2019

Environment Driven Underwater Camera-IMU Calibration for Monocular Visual-Inertial SLAM

ICRA 2019poster

Most state-of-the-art underwater vision systems are calibrated manually in shallow water and used in open seas without changing. However, the refractivity of the water is adaptively changed depending on the salinity, temperature, depth or other underwater environmental indexes, which inevitably gene…

Cited by 36SourceScholar
2019

Semantic-Transferable Weakly-Supervised Endoscopic Lesions Segmentation

ICCV 2019accepted

Weakly-supervised learning under image-level labels supervision has been widely applied to semantic segmentation of medical lesions regions. However, 1) most existing models rely on effective constraints to explore the internal representation of lesions, which only produces inaccurate and coarse les…

Cited by 61SourcePDFScholar
2017

Deep learning of directional truncated signed distance function for robust 3D object recognition

IROS 2017poster

In this paper, we develop a novel 3D object recognition algorithm to perform detection and pose estimation jointly. We focus on analyzing the advantages of the 3D point cloud relative to the RGB-D image and try to eliminate the unpredictability of output values that inevitably occurs in regression t…

Cited by 14SourceScholar
2016

A design of phase-closed-loop nanomachining control based ultrasonic vibration-assisted AFM

IROS 2016poster

This paper proposed a phase-closed-loop nanomachining control method to realize the directly control of machining depth based on ultrasonic vibration-assisted AFM. By using applied force to control the machining depth, conventional AFM machining approaches unable to machining a nanostructure with sp…

Cited by 0SourceScholar