← Search

Gan Sun

17 accepted papers

2026

PixelVLA: Advancing Pixel-level Understanding in Vision-Language-Action Model

ICLR 2026poster

Vision-Language-Action models (VLAs) are emerging as powerful tools for learning generalizable visuomotor control policies. However, current VLAs are mostly trained on large-scale image–text–action data and remain limited in two key ways: (i) they struggle with pixel-level scene understanding, and (…

Cited by 0SourceScholar
2026

R2-LIO: Real-Time and Robust LiDAR-Inertial Odometry in Dynamic Environments

ICRA 2026poster

LiDAR-Inertial Odometry (LIO) is crucial for robot navigation and autonomous driving. Most existing methods rely on the assumption of a static environment, indiscriminately using all LiDAR measurements for localization. However, LiDAR data acquired in urban scenes often contain dynamic objects such …

Cited by 0Scholar
2025

GLAM: Global-Local Variation Awareness in Mamba-based World Model

AAAI 2025technical

Mimicking the real interaction trajectory in the inference of the world model has been shown to improve the sample efficiency of model-based reinforcement learning (MBRL) algorithms. Many methods directly use known state sequences for reasoning. However, this approach fails to enhance the quality of…

2024

Cs2K: Class-specific and Class-shared Knowledge Guidance for Incremental Semantic Segmentation

ECCV 2024poster

"Incremental semantic segmentation endeavors to segment newly encountered classes while maintaining knowledge of old classes. However, existing methods either 1) lack guidance from class-specific knowledge (i.e., old class prototypes), leading to a bias towards new classes, or 2) constrain class-sha…

Cited by 1SourcePDFScholar
2024

GroupTrack: Multi-Object Tracking by Using Group Motion Patterns

IROS 2024poster

The main challenge of Multi-Object Tracking (MOT) lies in maintaining a distinctive identity for each target in dense crowds or occluded scenarios. Although the existing methods have achieved significantly progress by using robust object detectors or complex association strategies, they cannot effec…

Cited by 0SourceScholar
2024

Novel Object Synthesis via Adaptive Text-Image Harmony

NeurIPS 2024poster

In this paper, we study an object synthesis task that combines an object text with an object image to create a new object image. However, most diffusion models struggle with this task, \textit{i.e.}, often generating an object that predominantly reflects either the text or the image due to an imbala…

2022

Class-Incremental Gesture Recognition Learning with Out-of-Distribution Detection

IROS 2022poster

Gesture recognition is a popular human-computer interaction technology, which has been widely applied in many fields (e.g., autonomous driving, medical care, VR and AR). However, 1) most existing gesture recognition methods focus on the fixed recognition scenarios with several gestures, which could…

Cited by 13SourceScholar
2021

Confident Anchor-Induced Multi-Source Free Domain Adaptation

NeurIPS 2021poster

Unsupervised domain adaptation has attracted appealing academic attentions by transferring knowledge from labeled source domain to unlabeled target domain. However, most existing methods assume the source data are drawn from a single domain, which cannot be successfully applied to explore complement…

2021

Generative Partial Visual-Tactile Fused Object Clustering

AAAI 2021technical

Visual-tactile fused sensing for object clustering has achieved significant progresses recently, since the involvement of tactile modality can effectively improve clustering performance. However, the missing data (i.e., partial data) issues always happen due to occlusion and noises during the data c…

Cited by 17SourcePDFScholar
2020

CSCL: Critical Semantic-Consistent Learning for Unsupervised Domain Adaptation

ECCV 2020poster

Unsupervised domain adaptation without consuming annotation process for unlabeled target data attracts appealing interests in semantic segmentation. However, 1) existing methods neglect that not all semantic representations across domains are transferable, which cripples domain-wise transfer with un…

Cited by 56SourcePDFScholar
2020

What Can Be Transferred: Unsupervised Domain Adaptation for Endoscopic Lesions Segmentation

CVPR 2020poster

Unsupervised domain adaptation has attracted growing research attention on semantic segmentation. However, 1) most existing models cannot be directly applied into lesions transfer of medical images, due to the diverse appearances of same lesion among different datasets; 2) equal attention has been p…

Cited by 173PDFScholar
2019

Environment Driven Underwater Camera-IMU Calibration for Monocular Visual-Inertial SLAM

ICRA 2019poster

Most state-of-the-art underwater vision systems are calibrated manually in shallow water and used in open seas without changing. However, the refractivity of the water is adaptively changed depending on the salinity, temperature, depth or other underwater environmental indexes, which inevitably gene…

Cited by 36SourceScholar
2019

Semantic-Transferable Weakly-Supervised Endoscopic Lesions Segmentation

ICCV 2019accepted

Weakly-supervised learning under image-level labels supervision has been widely applied to semantic segmentation of medical lesions regions. However, 1) most existing models rely on effective constraints to explore the internal representation of lesions, which only produces inaccurate and coarse les…

Cited by 61SourcePDFScholar