← Search

Yanjiang Guo

16 accepted papers

2026

BagelVLA: Enhancing Long-Horizon Manipulation via Interleaved Vision-Language-Action Generation

RSS 2026poster

Equipping embodied agents with the ability to reason about tasks, foresee physical outcomes, and generate precise actions is essential for general-purpose manipulation. While recent Vision-Language-Action (VLA) models have leveraged pre-trained foundation models, they typically focus on either lingu…

Cited by 0SourceScholar
2026

Ctrl-World: A Controllable Generative World Model for Robot Manipulation

ICLR 2026poster

Generalist robot policies can now perform a wide range of manipulation skills, but evaluating and improving their ability with unfamiliar objects and instructions remains a significant challenge. Rigorous evaluation requires a large number of real-world rollouts, while systematic improvement demands…

Cited by 0SourcecodeScholar
2026

UniCoD: Enhancing Robot Policy via Unified Continuous and Discrete Representation Learning

ICML 2026poster

Building generalist robot policies that can handle diverse tasks in open-ended environments is a central challenge in robotics. To leverage knowledge from large-scale pretraining, prior work (VLA) has typically built generalist policies either on top of vision-language models (VLMs) or generative mo…

Cited by 0SourceScholar
2026

VLAW: Iterative Co-Improvement of Vision-Language-Action Policy and World Model

ICML 2026poster

The goal of this paper is to improve the performance and reliability of vision-language-action (VLA) models through iterative online interaction. Since collecting policy rollouts in the real world is expensive, we investigate whether a learned simulator—specifically, an action-conditioned video gene…

Cited by 0SourceScholar
2026

VLM4VLA: Revisiting Vision-Language-Models in Vision-Language-Action Models

ICLR 2026poster

Vision-Language-Action (VLA) models, which integrate pretrained large Vision-Language Models (VLMs) into their policy backbone, are gaining significant attention for their promising generalization capabilities. This paper revisits a fundamental yet seldom systematically studied question: how the cho…

Cited by 0SourceScholar
2026

villa-X: Enhancing Latent Action Modeling in Vision-Language-Action Models

ICLR 2026poster

Vision-Language-Action (VLA) models have emerged as a popular paradigm for learning robot manipulation policies that can follow language instructions and generalize to novel scenarios. Recent works have begun to explore the incorporation of latent actions, abstract representations of motion between…

Cited by 0SourcecodeScholar
2025

Improving Vision-Language-Action Model with Online Reinforcement Learning

ICRA 2025

Recent studies have successfully integrated large vision-language models (VLMs) into low-level robotic control by supervised fine-tuning (SFT) with expert robotic datasets, resulting in what we term vision-language-action (VLA) models. Although the VLA models are powerful, how to improve these large

Cited by 78SourceScholar
2025

UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent

ICML 2025poster

Recent advancements in Vision-Language-Action (VLA) models have leveraged pre-trained Vision-Language Models (VLMs) to improve the generalization capabilities. VLMs, typically pre-trained on vision-language understanding tasks, provide rich semantic knowledge and reasoning abilities. However, prior…

Cited by 2SourcePDFScholar
2025

Video Prediction Policy: A Generalist Robot Policy with Predictive Visual Representations

ICML 2025spotlight

Visual representations play a crucial role in developing generalist robotic policies. Previous vision encoders, typically pre-trained with single-image reconstruction or two-image contrastive learning, tend to capture static information, often neglecting the dynamic aspects vital for embodied tasks.…

2024

Advancing Humanoid Locomotion: Mastering Challenging Terrains with Denoising World Model Learning

RSS 2024poster

Humanoid robots, with their human-like skeletal structure, are especially suited for tasks in human-centric environments. However, this structure is accompanied by additional challenges in locomotion controller design, especially in complex real-world environments. As a result, existing humanoid rob…

2024

DoReMi: Grounding Language Model by Detecting and Recovering from Plan-Execution Misalignment

IROS 2024poster

Large language models (LLMs) encode a vast amount of semantic knowledge and possess remarkable understanding and reasoning capabilities. Previous work has explored how to ground LLMs in robotic tasks to generate feasible and executable textual plans. However, low-level execution in the physical worl…

Cited by 37SourceScholar
2024

HiRT: Enhancing Robotic Control with Hierarchical Robot Transformers

CoRL 2024poster

Large Vision-Language-Action (VLA) models, leveraging powerful pre-trained Vision-Language Models (VLMs) backends, have shown promise in robotic control due to their impressive generalization ability. However, the success comes at a cost. Their reliance on VLM backends with billions of parameters le…

Cited by 8SourceScholar
2024

Prediction with Action: Visual Policy Learning via Joint Denoising Process

NeurIPS 2024poster

Diffusion models have demonstrated remarkable capabilities in image generation tasks, including image editing and video creation, representing a good understanding of the physical world. On the other line, diffusion models have also shown promise in robotic control tasks by denoising actions, known…

Cited by 4SourcePDFScholar
2023

Decentralized Motor Skill Learning for Complex Robotic Systems

RA-L 2023

Reinforcement learning (RL) has achieved remarkable success in complex robotic systems (eg. quadruped locomotion). In previous works, the RL-based controller was typically implemented as a single neural network with concatenated observation input. However, the corresponding learned policy is highly

Cited by 9SourceScholar
2023

Zero-Shot Policy Transfer with Disentangled Task Representation of Meta-Reinforcement Learning

ICRA 2023poster

Humans are capable of abstracting various tasks as different combinations of multiple attributes. This perspective of compositionality is vital for human rapid learning and adaption since previous experiences from related tasks can be combined to generalize across novel compositional settings. In th…

Cited by 14SourceScholar
2022

Reinforcement learning with Demonstrations from Mismatched Task under Sparse Reward

CoRL 2022poster

Reinforcement learning often suffer from the sparse reward issue in real-world robotics problems. Learning from demonstration (LfD) is an effective way to eliminate this problem, which leverages collected expert data to aid online learning. Prior works often assume that the learning agent and the ex…

Cited by 6SourceScholar