← Search

Guangyan Chen

14 accepted papers

2026

Learning a Unified Latent Action Space from Videos with Action-centric Cycle Consistency

CVPR 2026

Video data provides a rich source beyond expensive action-labeled data for advancing robot learning. Recent approaches have demonstrated promising potential in leveraging video data by learning latent actions for policy training. The latent action tokenizer encodes latent actions between successive

Cited by 0SourceScholar
2025

GraphMimic: Graph-to-Graphs Generative Modeling from Videos for Policy Learning

CVPR 2025poster

Learning from demonstration is a powerful method for robotic skill acquisition. However, the significant expense of collecting such action-labeled robot data presents a major bottleneck. Video data, a rich data source encompassing diverse behavioral and physical knowledge, emerges as a promising alt…

Cited by 0SourcePDFScholar
2025

Hierarchical Autoregressive Modeling With Multi-Scale Refinement for Robot Policy Learning

RA-L 2025

While autoregressive models demonstrate remarkable success in text and image generation, their application to robot policies suffers from weak holistic comprehension, cumulative errors, and limited multi-modal modeling capabilities, particularly in long-horizon tasks or multi-modal scenarios. This p

Cited by 1SourceScholar
2025

High-Precision Object Pose Estimation Using Visual-Tactile Information for Dynamic Interactions in Robotic Grasping

ICRA 2025

In various robotic applications, understanding accurate object poses for robots is essential for high-precision tasks such as factory assembly or daily insertions. Tactile sensing, which compensates for visual information, offers rich texture-based or force-based data for object pose estimation. How

Cited by 0SourceScholar
2025

Human Demonstrations are Generalizable Knowledge for Robots

IROS 2025

Learning from human demonstrations is an emerging trend for designing intelligent robotic systems. However, previous methods typically regard videos as instructions, simply dividing videos into action sequences for robotic repetition, which pose obstacles to generalization to diverse tasks or object

Cited by 11SourceScholar
2025

ORA-NET: Enhancing Image Feature Matching through Oriented Overlapping Region Alignment

IROS 2025

Image feature matching is a fundamental task in computer vision. Existing local feature matching methods can establish robust correspondences between image pairs. However, these methods heavily rely on dense local image features, making them susceptible to significant perspective differences, charac

Cited by 0SourceScholar
2025

PartGrasp: Generalizable Part-level Grasping via Semantic-Geometric Alignment

IROS 2025

The ability to perform generalizable and precise grasping on functional object parts is a prerequisite for robotic manipulation in open environments. Recent foundation models have demonstrated promising semantic correspondence capabilities in guiding robots to grasp similar parts across objects with

Cited by 0SourcecodeScholar
2025

STEP Planner: Constructing cross-hierarchical subgoal tree as an embodied long-horizon task planner

IROS 2025

The ability to perform reliable long-horizon task planning is crucial for deploying robots in real-world environments. However, directly employing Large Language Models (LLMs) as action sequence generators often results in low success rates due to their limited reasoning ability for long-horizon emb

Cited by 4SourceScholar
2024

Fast and Robust Point Cloud Registration with Tree-based Transformer

ICRA 2024poster

Point cloud registration is essential in computer vision and robotics. Recently, transformer-based methods have achieved advanced point cloud registration performance. However, the standard attention mechanism utilized in these methods considers many low-relevance points, and it has difficulty focus…

Cited by 2SourcecodeScholar
2024

OpenGraph: Open-Vocabulary Hierarchical 3D Graph Representation in Large-Scale Outdoor Environments

RA-L 2024

Environment representations endowed with sophisticated semantics are pivotal for facilitating seamless interaction between robots and humans, enabling them to effectively carry out various tasks. Open-vocabulary representation, powered by Visual-Language models (VLMs), possesses inherent advantages,

Cited by 38SourcecodeScholar
2024

VLMimic: Vision Language Models are Visual Imitation Learner for Fine-grained Actions

NeurIPS 2024poster

Visual imitation learning (VIL) provides an efficient and intuitive strategy for robotic systems to acquire novel skills. Recent advancements in Vision Language Models (VLMs) have demonstrated remarkable performance in vision and language reasoning capabilities for VIL tasks. Despite the progress, c…

Cited by 5SourcePDFScholar
2023

Deep Interactive Full Transformer Framework for Point Cloud Registration

ICRA 2023poster

Point cloud registration is a crucial technology in the fields of robotics and computer vision. Despite the significant advances in point cloud registration enabled by Transformer-based methods, limitations persist due to indistinct feature extraction, noise sensitivity, and outlier handling. These…

Cited by 7SourcecodeScholar
2023

PointGPT: Auto-regressively Generative Pre-training from Point Clouds

NeurIPS 2023poster

Large language models (LLMs) based on the generative pre-training transformer (GPT) have demonstrated remarkable effectiveness across a diverse range of downstream tasks. Inspired by the advancements of the GPT, we present PointGPT, a novel approach that extends the concept of GPT to point clouds, a…

2023

Rethinking Point Cloud Registration as Masking and Reconstruction

ICCV 2023poster

Point cloud registration is essential in computer vision and robotics. In this paper, a critical observation is made that the invisible parts of each point cloud can be directly utilized as inherent masks, and the aligned point cloud pair can be regarded as the reconstruction target. Motivated by th…

Cited by 13PDFcodeScholar