← Search

Yutong Hu

11 accepted papers

2026

AR-VLA: Autoregressive Action Expert for Vision–Language–Action Models

RSS 2026poster

We propose a standalone autoregressive (AR) Action Expert that generates actions as a continuous causal sequence while conditioning on refreshable vision-language prefixes. In contrast to existing Vision-Language-Action (VLA) models and diffusion policies that reset temporal context with each new ob…

Cited by 0SourceScholar
2026

Mini Diffuser: Fast Multi-Task Diffusion Policy Training Using Two-Level Mini-Batches

RA-L 2026

We present a method that reduces, by an order of magnitude, the time and memory needed to train multi-task vision-language robotic diffusion policies. This improvement arises from a previously underexplored distinction between action diffusion and the image diffusion techniques that inspired it: In

Cited by 0SourcecodeScholar
2026

Mini Diffuser: Fast Multi-Task Diffusion Policy Training Using Two-Level Mini-Batches

ICRA 2026poster

We present a method that reduces, by an order of magnitude, the time and memory needed to train multi-task vision-language robotic diffusion policies. This improvement arises from a previously underexplored distinction between action diffusion and the image diffusion techniques that inspired it: In …

2025

M^3PC: Test-time Model Predictive Control using Pretrained Masked Trajectory Model

ICLR 2025poster

Recent work in Offline Reinforcement Learning (RL) has shown that a unified transformer trained under a masked auto-encoding objective can effectively capture the relationships between different modalities (e.g., states, actions, rewards) within given trajectory datasets. However, this information…

2024

ELLA: Empowering LLMs for Interpretable, Accurate and Informative Legal Advice

ACL 2024system demonstrations

Despite remarkable performance in legal consultation exhibited by legal Large Language Models(LLMs) combined with legal article retrieval components, there are still cases when the advice given is incorrect or baseless. To alleviate these problems, we propose ELLA, a tool for Empowering LLMs for int…

2023

Align-then-Enhance: Multilingual Entailment Graph Enhancement with Soft Predicate Alignment

ACL 2023findings

Entailment graphs (EGs) with predicates as nodes and entailment relations as edges are typically incomplete, while EGs in different languages are often complementary to each other. In this paper, we propose a new task, multilingual entailment graph enhancement, which aims to utilize the entailment i…

Cited by 3SourcePDFScholar
2023

More than Classification: A Unified Framework for Event Temporal Relation Extraction

ACL 2023long

Event temporal relation extraction (ETRE) is usually formulated as a multi-label classification task, where each type of relation is simply treated as a one-hot label. This formulation ignores the meaning of relations and wipes out their intrinsic dependency. After examining the relation definitions…

2022

SO-SLAM: Semantic Object SLAM With Scale Proportional and Symmetrical Texture Constraints

RA-L 2022

Object SLAM introduces the concept of objects into Simultaneous Localization and Mapping (SLAM) and helps understand indoor scenes for mobile robots and object-level interactive applications. The state-of-art object SLAM systems face challenges such as partial observations, occlusions, unobservable

Cited by 81SourcecodeScholar