← Search

Bingqian Lin

13 accepted papers

2025

Affordances-Oriented Planning Using Foundation Models for Continuous Vision-Language Navigation

AAAI 2025technical

LLM-based agents have demonstrated impressive zero-shot performance in vision-language navigation (VLN) task. However, existing LLM-based methods often focus only on solving high-level task planning by selecting nodes in predefined navigation graphs for movements, overlooking low-level control in na…

Cited by 8SourcePDFScholar
2025

MineAnyBuild: Benchmarking Spatial Planning for Open-world AI Agents

NeurIPS 2025poster

Spatial Planning is a crucial part in the field of spatial intelligence, which requires the understanding and planning about object arrangements in space perspective. AI agents with the spatial planning ability can better adapt to various real-world applications, including robotic manipulation, auto…

Cited by 0SourcecodeScholar
2025

PhyBlock: A Progressive Benchmark for Physical Understanding and Planning via 3D Block Assembly

NeurIPS 2025poster

While vision-language models (VLMs) have demonstrated promising capabilities in reasoning and planning for embodied agents, their ability to comprehend physical phenomena, particularly within structured 3D environments, remains severely limited. To close this gap, we introduce PhyBlock, a progressiv…

Cited by 0SourceScholar
2025

Structured Preference Optimization for Vision-Language Long-Horizon Task Planning

EMNLP 2025

Existing vision-language planning methods perform well on short-horizon tasks but struggle with long-horizon reasoning in dynamic environments due to the difficulty of training models to generate high-quality reasoning processes. To address this, we propose Structured Preference Optimization (SPO),

Cited by 0SourcePDFScholar
2024

MapGPT: Map-Guided Prompting with Adaptive Path Planning for Vision-and-Language Navigation

ACL 2024long

Embodied agents equipped with GPT as their brain have exhibited extraordinary decision-making and generalization abilities across various tasks. However, existing zero-shot agents for vision-and-language navigation (VLN) only prompt the GPT-4 to select potential locations within localized environmen…

Cited by 30SourcePDFScholar
2024

PIVOT-R: Primitive-Driven Waypoint-Aware World Model for Robotic Manipulation

NeurIPS 2024poster

Language-guided robotic manipulation is a challenging task that requires an embodied agent to follow abstract user instructions to accomplish various complex manipulation tasks. Previous work generally maps instructions and visual perceptions directly to low-level executable actions, neglecting the…

Cited by 1SourcePDFScholar
2023

Actional Atomic-Concept Learning for Demystifying Vision-Language Navigation

AAAI 2023technical

Vision-Language Navigation (VLN) is a challenging task which requires an agent to align complex visual observations to language instructions to reach the goal position. Most existing VLN agents directly learn to align the raw directional features and visual features trained using one-hot labels to l…

Cited by 5SourcePDFScholar
2023

Dynamic Graph Enhanced Contrastive Learning for Chest X-Ray Report Generation

CVPR 2023poster

Automatic radiology reporting has great clinical potential to relieve radiologists from heavy workloads and improve diagnosis interpretation. Recently, researchers have enhanced data-driven neural networks with medical knowledge graphs to eliminate the severe visual and textual bias in this task. Th…

2022

ADAPT: Vision-Language Navigation With Modality-Aligned Action Prompts

CVPR 2022poster

Vision-Language Navigation (VLN) is a challenging task that requires an embodied agent to perform action-level modality alignment, i.e., make instruction-asked actions sequentially in complex visual environments. Most existing VLN agents learn the instruction-path data directly and cannot sufficient…

Cited by 58PDFScholar
2022

Contrastive Instruction-Trajectory Learning for Vision-Language Navigation

AAAI 2022technical

The vision-language navigation (VLN) task requires an agent to reach a target with the guidance of natural language instruction. Previous works learn to navigate step-by-step following an instruction. However, these works may fail to discriminate the similarities and discrepancies across instruction…

2022

RelCLIP: Adapting Language-Image Pretraining for Visual Relationship Detection via Relational Contrastive Learning

EMNLP 2022main

Conventional visual relationship detection models only use the numeric ids of relation labels for training, but ignore the semantic correlation between the labels, which leads to severe training biases and harms the generalization ability of representations. In this paper, we introduce compact langu…

2020

Vision-Dialog Navigation by Exploring Cross-Modal Memory

CVPR 2020poster

Vision-dialog navigation posed as a new holy-grail task in vision-language disciplinary targets at learning an agent endowed with the capability of constant conversation for help with natural language and navigating according to human responses. Besides the common challenges faced in visual language…

Cited by 55PDFcodeScholar