← Search

Sitong Mao

10 accepted papers

2026

Discrete Diffusion VLA: Bringing Discrete Diffusion to Action Decoding in Vision-Language-Action Policies

ICML 2026poster

Vision–Language–Action (VLA) models adapt large vision–language backbones to map images and instructions into robot actions. However, prevailing VLAs either generate actions autoregressively in a fixed left-to-right order or attach separate diffusion heads outside the backbone, fragmenting informati…

Cited by 0SourceScholar
2026

Spatial Reasoning with Vision-Language Models in Ego-Centric Multi-View Scenes

ICLR 2026poster

Understanding 3D spatial relationships remains a major limitation of current Vision-Language Models (VLMs). Prior work has addressed this issue by creating spatial question-answering (QA) datasets based on single images or indoor videos. However, real-world embodied AI agents—such as robots and self…

Cited by 0SourcecodeScholar
2025

ASCENT: Autonomous Skill Learning Toward Complex Embodied Tasks With Foundation Models

ICRA 2025

Collecting data from simulated scenarios for training robotic skills provides a safer and more controllable alternative to real-world environments. However, it demands considerable effort, including the manual construction of simulation environments, the careful design of tasks, and the challenge of

Cited by 0SourceScholar
2025

Ms. NAMI: Multimodal Semantic Navigation on Relative Metric Intention Graph

ICRA 2025

Embodied navigation in unknown environments presents the significant challenge of integrating tasks with multimodal goals into a unified framework. In this paper, we propose the Multimodal Semantic Navigation on Relative Metric Intention Graph (Ms. NAMI), a framework that integrates various navigati

Cited by 0SourceScholar
2025

PanopticSplatting: End-to-End Panoptic Gaussian Splatting

IROS 2025

Open-vocabulary panoptic reconstruction is a challenging task for simultaneous scene reconstruction and understanding. Recently, methods have been proposed for 3D scene understanding based on Gaussian splatting. However, these methods are multi-staged, suffering from the accumulated errors and the d

Cited by 2SourceScholar
2024

Let Occ Flow: Self-Supervised 3D Occupancy Flow Prediction

CoRL 2024poster

Accurate perception of the dynamic environment is a fundamental task for autonomous driving and robot systems. This paper introduces Let Occ Flow, the first self-supervised work for joint 3D occupancy and occupancy flow prediction using only camera inputs, eliminating the need for 3D annotations. Ut…

Cited by 10SourceScholar
2024

PanopticRecon: Leverage Open-vocabulary Instance Segmentation for Zero-shot Panoptic Reconstruction

IROS 2024

Panoptic reconstruction is a challenging task in 3D scene understanding. However, most existing methods heavily rely on pre-trained semantic segmentation models and known 3D object bounding boxes for 3D panoptic segmentation, which is not available for in-the-wild scenes. In this paper, we propose a

Cited by 8SourceScholar
2024

SCALE: Self-Correcting Visual Navigation for Mobile Robots via Anti-Novelty Estimation

ICRA 2024poster

Although visual navigation has been extensively studied using deep reinforcement learning, online learning for real-world robots remains a challenging task. Recent work directly learned from offline dataset to achieve broader generalization in the real-world tasks, which, however, faces the out-of-d…

Cited by 2SourcecodeScholar
2024

Scale Disparity of Instances in Interactive Point Cloud Segmentation

IROS 2024poster

Interactive point cloud segmentation has become a pivotal task for understanding 3D scenes, enabling users to guide segmentation models with simple interactions such as clicks, therefore significantly reducing the effort required to tailor models to diverse scenarios and new categories. However, in…

Cited by 2SourceScholar
2023

NF-Atlas: Multi-Volume Neural Feature Fields for Large Scale LiDAR Mapping

RA-L 2023

LiDAR Mapping has been a long-standing problem in robotics. Recent progress in neural implicit representation has brought new opportunities to robotic mapping. In this letter, we propose the multi-volume neural feature fields, called NF-Atlas, which bridge the neural feature volumes with pose graph

Cited by 21SourceScholar