← Search

Zihan Lan

3 accepted papers

2026

BFA++: Hierarchical Best-Feature-Aware Token Prune for Multi-View Vision Language Action Model

RA-L 2026

Vision-Language-Action (VLA) models have achieved significant breakthroughs by leveraging Large Vision Language Models (VLMs) to jointly interpret instructions and visual inputs. However, the substantial increase in visual tokens, particularly from multi-view inputs, poses serious challenges to real

Cited by 0SourceScholar
2026

BFA: Best-Feature-Aware Fusion for Multi-View Fine-Grained Manipulation

ICRA 2026poster

In real-world scenarios, multi-view cameras are typically employed for fine-grained manipulation tasks. Existing approaches (e.g., ACT ) tend to treat multi-view features equally and directly concatenate them for policy learning. How ever, it will introduce redundant visual information and bring hig…

2025

BFA: Best-Feature-Aware Fusion for Multi-View Fine-Grained Manipulation

RA-L 2025

In real-world scenarios, multi-view cameras are typically employed for fine-grained manipulation tasks. Existing approaches (e.g., ACT [1]) tend to treat multi-view features equally and directly concatenate them for policy learning. However, it will introduce redundant visual information and bring h

Cited by 8SourceScholar