← Search

Juexiao Zhang

8 accepted papers

2026

CRAG: Can 3D Generative Models Help 3D Assembly?

ICML 2026poster

Most existing 3D assembly methods treat the problem as pure pose estimation, rearranging observed parts via rigid transformations. In contrast, human assembly naturally couples structural reasoning with holistic shape inference. Inspired by this intuition, we reformulate 3D assembly as a joint probl…

Cited by 0SourceScholar
2025

CityWalker: Learning Embodied Urban Navigation from Web-Scale Videos

CVPR 2025poster

Navigating dynamic urban environments presents significant challenges for embodied agents, requiring advanced spatial reasoning and adherence to common-sense norms. Despite progress, existing visual navigation methods struggle in map-free or off-street settings, limiting the deployment of autonomous…

2025

URLOST: Unsupervised Representation Learning without Stationarity or Topology

ICLR 2025poster

Unsupervised representation learning has seen tremendous progress. However, it is constrained by its reliance on domain specific stationarity and topology, a limitation not found in biological intelligence systems. For instance, unlike computer vision, human vision can process visual signals sampled…

Cited by 1SourcePDFScholar
2025

VLM See, Robot Do: Human Demo Video to Robot Action Plan via Vision Language Model

IROS 2025

Large Vision Language Models (VLMs) have been adopted in robotics for their strong common sense understanding and generalization capabilities. Existing works leverage VLMs for task and motion planning based on language instructions and robot observations. In this work, we explore using VLM to interp

Cited by 45SourcecodeScholar
2024

LUWA Dataset: Learning Lithic Use-Wear Analysis on Microscopic Images

CVPR 2024highlight

Lithic Use-Wear Analysis (LUWA) using microscopic images is an underexplored vision-for-science research area. It seeks to distinguish the worked material which is critical for understanding archaeological artifacts material interactions tool functionalities and dental records. However this challeng…

Cited by 4SourcePDFScholar
2022

Multi-Robot Scene Completion: Towards Task-Agnostic Collaborative Perception

CoRL 2022poster

Collaborative perception learns how to share information among multiple robots to perceive the environment better than individually done. Past research on this has been task-specific, such as detection or segmentation. Yet this leads to different information sharing for different tasks, hindering th…

Cited by 55SourcecodeScholar