← Search

Zhibin Tang

5 accepted papers

2025

ChatVLA-2: Vision-Language-Action Model with Open-World Reasoning

NeurIPS 2025poster

Vision-language-action (VLA) models have emerged as the next generation of models in robotics. However, despite leveraging powerful pre-trained Vision-Language Models (VLMs), existing end-to-end VLA systems often lose key capabilities during fine-tuning as the model adapts to specific robotic tasks.…

Cited by 0SourceScholar
2025

CoA-VLA: Improving Vision-Language-Action Models via Visual-Text Chain-of-Affordance

ICCV 2025poster

Robot foundation models, particularly Vision-Language-Action (VLA) models, have garnered significant attention for their ability to enhance robot policy learning, greatly improving robot's generalization and robustness. OpenAI's recent model, O1, showcased impressive capabilities in solving complex…

Cited by 0SourcePDFScholar
2025

DexVLA: Vision-Language Model with Plug-In Diffusion Expert for General Robot Control

CoRL 2025poster

Enabling robots to perform diverse tasks across varied environments is a central challenge in robot learning. While vision-language-action (VLA) models have shown promise for generalizable robot skills, realizing their full potential requires addressing limitations in action representation and effic…

Cited by 0SourceScholar
2025

DiffusionVLA: Scaling Robot Foundation Models via Unified Diffusion and Autoregression

ICML 2025poster

In this paper, we present DiffusionVLA, a novel framework that integrates autoregressive reasoning with diffusion policies to address the limitations of existing methods: while autoregressive Vision-Language-Action (VLA) models lack precise and robust action generation, diffusion-based policies inhe…

Cited by 0SourcePDFScholar
2025

TinyVLA: Toward Fast, Data-Efficient Vision-Language-Action Models for Robotic Manipulation

RA-L 2025

Vision-Language-Action (VLA) models have shown remarkable potential in visuomotor control and instruction comprehension through end-to-end learning processes. However, current VLA models face significant challenges: they are slow during inference and require extensive pre-training on large amounts o

Cited by 303SourceScholar