← Search

Te Cui

9 accepted papers

2026

Learning a Unified Latent Action Space from Videos with Action-centric Cycle Consistency

CVPR 2026

Video data provides a rich source beyond expensive action-labeled data for advancing robot learning. Recent approaches have demonstrated promising potential in leveraging video data by learning latent actions for policy training. The latent action tokenizer encodes latent actions between successive

Cited by 0SourceScholar
2025

GraphMimic: Graph-to-Graphs Generative Modeling from Videos for Policy Learning

CVPR 2025poster

Learning from demonstration is a powerful method for robotic skill acquisition. However, the significant expense of collecting such action-labeled robot data presents a major bottleneck. Video data, a rich data source encompassing diverse behavioral and physical knowledge, emerges as a promising alt…

Cited by 0SourcePDFScholar
2025

Hierarchical Autoregressive Modeling With Multi-Scale Refinement for Robot Policy Learning

RA-L 2025

While autoregressive models demonstrate remarkable success in text and image generation, their application to robot policies suffers from weak holistic comprehension, cumulative errors, and limited multi-modal modeling capabilities, particularly in long-horizon tasks or multi-modal scenarios. This p

Cited by 1SourceScholar
2025

High-Precision Object Pose Estimation Using Visual-Tactile Information for Dynamic Interactions in Robotic Grasping

ICRA 2025

In various robotic applications, understanding accurate object poses for robots is essential for high-precision tasks such as factory assembly or daily insertions. Tactile sensing, which compensates for visual information, offers rich texture-based or force-based data for object pose estimation. How

Cited by 0SourceScholar
2025

Human Demonstrations are Generalizable Knowledge for Robots

IROS 2025

Learning from human demonstrations is an emerging trend for designing intelligent robotic systems. However, previous methods typically regard videos as instructions, simply dividing videos into action sequences for robotic repetition, which pose obstacles to generalization to diverse tasks or object

Cited by 11SourceScholar
2025

ORA-NET: Enhancing Image Feature Matching through Oriented Overlapping Region Alignment

IROS 2025

Image feature matching is a fundamental task in computer vision. Existing local feature matching methods can establish robust correspondences between image pairs. However, these methods heavily rely on dense local image features, making them susceptible to significant perspective differences, charac

Cited by 0SourceScholar
2025

TASeg: Text-aware RGB-T Semantic Segmentation based on Fine-tuning Vision Foundation Models

IROS 2025

Reliable semantic segmentation of open environments is essential for intelligent systems, yet significant problems remain: 1) Existing RGB-T semantic segmentation models mainly rely on low-level visual features and lack high-level textual information, which struggle with accurate segmentation when c

Cited by 1SourceScholar
2024

Robust Collaborative Perception against Temporal Information Disturbance

ICRA 2024poster

Collaborative perception facilitates a more comprehensive representation of the environment by leveraging complementary information shared among various agents and sensors. However, practical applications often encounter information disturbance which includes perception packet loss and time delays,…

Cited by 3SourcecodeScholar
2024

VLMimic: Vision Language Models are Visual Imitation Learner for Fine-grained Actions

NeurIPS 2024poster

Visual imitation learning (VIL) provides an efficient and intuitive strategy for robotic systems to acquire novel skills. Recent advancements in Vision Language Models (VLMs) have demonstrated remarkable performance in vision and language reasoning capabilities for VIL tasks. Despite the progress, c…

Cited by 5SourcePDFScholar