← Search

Shengxiang Sun

3 accepted papers

2026

Manual2Skill++: Connector-Aware General Robotic Assembly from Instruction Manuals Via Vision–Language Models

ICRA 2026poster

Assembly hinges on reliably forming connections between parts; yet most robotic approaches plan assembly sequences and part poses while treating connectors as an afterthought. Connections represent the foundational physical constraints of assembly execution; while task planning sequences operations,…

2025

Manual2Skill: Learning to Read Manuals and Acquire Robotic Skills for Furniture Assembly Using Vision-Language Models

RSS 2025poster

Humans possess an extraordinary ability to understand and execute complex manipulation tasks by interpreting abstract instruction manuals. For robots, however, this capability remains a substantial challenge, as they lack the ability to interpret abstract instructions and translate them into executa…

Cited by 1PDFcodeScholar
2025

SAFE: Multitask Failure Detection for Vision-Language-Action Models

NeurIPS 2025poster

While vision-language-action models (VLAs) have shown promising robotic behaviors across a diverse set of manipulation tasks, they achieve limited success rates when deployed on novel tasks out of the box. To allow these policies to safely interact with their environments, we need a failure detector…

Cited by 0SourcecodeScholar