← Search

Chengxuan Li

3 accepted papers

2026

Predicting What Matters: Robust Generalist Robot Policy Learning via Future Semantic Mask

ICML 2026poster

World models derived from large-scale video generative pre-training have emerged as a promising paradigm for generalist robot policy learning. However, standard approaches often focus on high-fidelity RGB video prediction, but this can result in overfitting to irrelevant factors, such as dynamic bac…

Cited by 0SourceScholar
2026

TTF-VLA: Temporal Token Fusion via Pixel-Attention Integration for Vision-Language-Action Models

AAAI 2026technical

Vision-Language-Action (VLA) models process visual inputs independently at each timestep, discarding valuable temporal information inherent in robotic manipulation tasks. This frame-by-frame processing makes models vulnerable to visual noise while ignoring the substantial coherence between consecuti

Cited by 0SourcePDFScholar