← Search

Tianyuan Yuan

9 accepted papers

2026

DepthVLA: Enhancing Vision-Language-Action Models with Depth-Aware Spatial Reasoning

ICRA 2026poster

Vision-Language-Action (VLA) models have recently shown impressive generalization and language-guided manipulation capabilities. However, their performance degrades on tasks requiring precise spatial reasoning due to limited spatial reasoning inherited from Vision-Language Models (VLMs). Existing VL…

2026

FASTer: Toward Powerful and Efficient Autoregressive Vision–Language–Action Models with Learnable Action Tokenizer and Block-wise Decoding

ICLR 2026poster

Autoregressive vision-language-action (VLA) models have recently demonstrated strong capabilities in robotic manipulation. However, their core process of action tokenization often involves a trade-off between reconstruction fidelity and inference efficiency. We introduce \textbf{FASTer}, a unified f…

Cited by 0SourceScholar
2025

Diffusion-Based Generative Models for 3D Occupancy Prediction in Autonomous Driving

ICRA 2025

Accurately predicting 3D occupancy grids from visual inputs is critical for autonomous driving, but current discriminative methods struggle with noisy data, incomplete observations, and the complex structures inherent in 3D scenes. In this work, we reframe 3D occupancy prediction as a generative mod

Cited by 5SourceScholar
2025

LONG3R: Long Sequence Streaming 3D Reconstruction

ICCV 2025poster

Recent advancements in multi-view scene reconstruction have been significant, yet existing methods face limitations when processing streams of input images. These methods either rely on time-consuming offline optimization or are restricted to shorter sequences, hindering their applicability in real-…

2024

P-MapNet: Far-Seeing Map Generator Enhanced by Both SDMap and HDMap Priors

RA-L 2024

Autonomous vehicles are gradually entering city roads today, with the help of high-definition maps (HDMaps). However, the reliance on HDMaps prevents autonomous vehicles from stepping into regions without this expensive digital infrastructure. This fact drives many researchers to study online HDMap

Cited by 62SourceScholar
2023

On Uni-Modal Feature Learning in Supervised Multi-Modal Learning

ICML 2023poster

We abstract the features (i.e. learned representations) of multi-modal data into 1) uni-modal features, which can be learned from uni-modal training, and 2) paired features, which can only be learned from cross-modal interactions. Multi-modal models are expected to benefit from cross-modal interacti…

2023

VectorMapNet: End-to-end Vectorized HD Map Learning

ICML 2023poster

Autonomous driving systems require High-Definition (HD) semantic maps to navigate around urban roads. Existing solutions approach the semantic mapping problem by offline manual annotation, which suffers from serious scalability issues. Recent learning-based methods produce dense rasterized segmentat…