← Search

Yuxing Long

10 accepted papers

2026

CorrectNav: Self-Correction Flywheel Empowers Vision-Language-Action Navigation Model

AAAI 2026technical

Existing vision-and-language navigation models often deviate from the correct trajectory when executing instructions. However, these models lack effective error correction capability, hindering their recovery from errors. To address this challenge, we propose Self-correction Flywheel, a novel post-t

Cited by 25SourcePDFScholar
2026

NavSpace: How Navigation Agents Follow Spatial Intelligence Instructions

ICRA 2026poster

Instruction-following navigation is a key step toward embodied intelligence. Prior benchmarks mainly focus on semantic understanding but overlook systematically evaluating navigation agents' spatial perception and reasoning capabilities. In this work, we introduce the NavSpace benchmark, which conta…

2026

RealAppiance: Let High-fidelity Appliance Assets Controllable and Workable as Aligned Real Manauls

CVPR 2026

Existing appliance assets suffer from poor rendering, incomplete mechanisms, and misalignment with manuals, leading to simulation-reality gaps that hinder appliance manipulation development. In this work, we introduce the RealAppliance dataset, comprising 100 high-fidelity appliances with complete p

Cited by 0SourceScholar
2025

CheckManual: A New Challenge and Benchmark for Manual-based Appliance Manipulation

CVPR 2025highlight

Correct use of electrical appliances has significantly improved human life quality. Unlike simple tools that can be manipulated with common sense, different parts of electrical appliances have specific functions defined by manufacturers. If we want the robot to heat bread by microwave, we should ena…

Cited by 0SourcePDFScholar
2024

Bridging Zero-shot Object Navigation and Foundation Models through Pixel-Guided Navigation Skill

ICRA 2024poster

Zero-shot object navigation is a challenging task for home-assistance robots. This task emphasizes visual grounding, commonsense inference and locomotion abilities, where the first two are inherent in foundation models. But for the locomotion part, most works still depend on map-based planning appro…

Cited by 39SourcecodeScholar
2024

Discuss Before Moving: Visual Language Navigation via Multi-expert Discussions

ICRA 2024poster

Visual language navigation (VLN) is an embodied task demanding a wide range of skills encompassing understanding, perception, and planning. For such a multifaceted challenge, previous VLN methods totally rely on one model’s own thinking to make predictions within one round. However, existing models,…

Cited by 51SourceScholar
2024

InstructNav: Zero-shot System for Generic Instruction Navigation in Unexplored Environment

CoRL 2024poster

Enabling robots to navigate following diverse language instructions in unexplored environments is an attractive goal for human-robot interaction. However, this goal is challenging because different navigation tasks require different strategies. The scarcity of instruction navigation data hinders tra…

Cited by 34SourceScholar
2024

ManipLLM: Embodied Multimodal Large Language Model for Object-Centric Robotic Manipulation

CVPR 2024poster

Robot manipulation relies on accurately predicting contact points and end-effector directions to ensure successful operation. However learning-based robot manipulation trained on a limited category within a simulator often struggles to achieve generalizability especially when confronted with extensi…

Cited by 54SourcePDFScholar
2023

Multimodal Recommendation Dialog with Subjective Preference: A New Challenge and Benchmark

ACL 2023findings

Existing multimodal task-oriented dialog data fails to demonstrate the diverse expressions of user subjective preferences and recommendation acts in the real-life shopping scenario. This paper introduces a new dataset SURE (Multimodal Recommendation Dialog with Subjective Preference), which contains…

2023

SPRING: Situated Conversation Agent Pretrained with Multimodal Questions from Incremental Layout Graph

AAAI 2023technical

Existing multimodal conversation agents have shown impressive abilities to locate absolute positions or retrieve attributes in simple scenarios, but they fail to perform well when complex relative positions and information alignments are involved, which poses a bottleneck in response quality. In thi…