← Search

Yuanliang Ju

3 accepted papers

2026

MomaGraph: State-Aware Unified Scene Graphs with Vision-Language Models for Embodied Task Planning

ICLR 2026oral

Mobile manipulators in households must both navigate and manipulate. This requires a compact, semantically rich scene representation that captures where objects are, how they function, and which parts are actionable. Scene graphs are a natural choice, yet prior work often separates spatial and funct…

Cited by 0SourcecodeScholar
2025

SAFE: Multitask Failure Detection for Vision-Language-Action Models

NeurIPS 2025poster

While vision-language-action models (VLAs) have shown promising robotic behaviors across a diverse set of manipulation tasks, they achieve limited success rates when deployed on novel tasks out of the box. To allow these policies to safely interact with their environments, we need a failure detector…

Cited by 0SourcecodeScholar
2024

ImOV3D: Learning Open Vocabulary Point Clouds 3D Object Detection from Only 2D Images

NeurIPS 2024poster

Open-vocabulary 3D object detection (OV-3Det) aims to generalize beyond the limited number of base categories labeled during the training phase. The biggest bottleneck is the scarcity of annotated 3D data, whereas 2D image datasets are abundant and richly annotated. Consequently, it is intuitive to…