← Search

Cong Tai

4 accepted papers

2026

RealMirror: A Comprehensive, Open-Source Vision-Language-Action Platform for Embodied AI

ICRA 2026poster

The emerging field of Vision-Language-Action (VLA) for humanoid robots faces several fundamental challenges, including the high cost of data acquisition, the lack of a standardized benchmark, and the significant gap between simulation and the real world. To overcome these obstacles, we propose RealM…

2024

FusedNet: End-to-End Mobile Robot Relocalization in Dynamic Large-Scale Scene

RA-L 2024

To improve robot relocalization accuracy in both static and dynamic environments, we introduce a novel network, FusedNet, which incorporates a cross-attention to fuse global and local image features for end-to-end relocalization. This approach relies solely on a monocular camera sensor that is fixed

Cited by 5SourceScholar
2024

GRID: Scene-Graph-based Instruction-driven Robotic Task Planning

IROS 2024

Recent works have shown that Large Language Models (LLMs) can facilitate the grounding of instructions for robotic task planning. Despite this progress, most existing works have primarily focused on utilizing raw images to aid LLMs in understanding environmental information. However, this approach n

Cited by 45SourcecodeScholar
2024

Mobile Robot Oriented Large-Scale Indoor Dataset for Dynamic Scene Understanding

ICRA 2024poster

Most existing robotic datasets capture static scene data and thus are limited in evaluating robots’ dynamic performance. To address this, we present a mobile robot oriented large-scale indoor dataset, denoted as THUD (Tsinghua University Dynamic) robotic dataset, for training and evaluating their dy…

Cited by 9SourceScholar