← Search

Xinyang Tong

6 accepted papers

2026

HiF-VLA: Hindsight, Insight and Foresight through Motion Representation for Vision-Language-Action Models

CVPR 2026

Vision-Language-Action (VLA) models have recently enabled robotic manipulation by grounding visual and linguistic cues into actions. However, most VLAs assume the Markov property, relying only on the current observation and thus suffering from temporal myopia that degrades long-horizon coherence. In

Cited by 0SourcecodeScholar
2026

VLA-Adapter: An Effective Paradigm for Tiny-Scale Vision-Language-Action Model

AAAI 2026technical

Vision-Language-Action (VLA) models typically bridge the gap between perceptual and action spaces by pre-training a large-scale Vision-Language Model (VLM) on robotic data. While this approach greatly enhances performance, it also incurs significant training costs. In this paper, we investigate how

Cited by 0SourcePDFScholar
2025

A2Seek: Towards Reasoning-Centric Benchmark for Aerial Anomaly Understanding

NeurIPS 2025poster

While unmanned aerial vehicles (UAVs) offer wide-area, high-altitude coverage for anomaly detection, they face challenges such as dynamic viewpoints, scale variations, and complex scenes. Existing datasets and methods, mainly designed for fixed ground-level views, struggle to adapt to these conditio…

Cited by 0SourcecodeScholar
2025

Long-VLA: Unleashing Long-Horizon Capability of Vision Language Action Model for Robot Manipulation

CoRL 2025poster

Vision-Language-Action (VLA) models have become a cornerstone in robotic policy learning, leveraging large-scale multimodal data for robust and scalable control. However, existing VLA frameworks primarily address short-horizon tasks, and their effectiveness on long-horizon, multi-step robotic manipu…

Cited by 0SourceScholar
2025

MoRE: Unlocking Scalability in Reinforcement Learning for Quadruped Vision-Language-Action Models

ICRA 2025

Developing versatile quadruped robots that can smoothly perform various actions and tasks in real-world environments remains a significant challenge. This paper introduces a novel vision-language-action (VLA) model, mixture of robotic experts (MoRE), for quadruped robots that aim to introduce reinfo

Cited by 24SourceScholar
2025

Quart-Online: Latency-Free Multimodal Large Language Model for Quadruped Robot Learning

ICRA 2025

This paper addresses the inherent inference latency challenges associated with deploying multimodal large language models (MLLM) in quadruped vision-language-action (QUAR-VLA) tasks. Our investigation reveals that conventional parameter reduction techniques ultimately impair the performance of the l

Cited by 1SourcecodeScholar