← Search

Zuojin Tang

1 accepted papers

2025

VLASCD: A Visual Language Action Model for Simultaneous Chatting and Decision Making

EMNLP 2025

Recent large pretrained models such as LLMs (e.g., GPT series) and VLAs (e.g., OpenVLA) have achieved notable progress on multimodal tasks, yet they are built upon a multi-input single-output (MISO) paradigm. We show that this paradigm fundamentally limits performance in multi-input multi-output (MI