← Search

Li Ji

5 accepted papers

2026

HiMe: Hierarchical Embodied Memory for Long-Horizon Vision-Language-Action Control

ICML 2026poster

Current Vision-Language-Action (VLA) models excel at robotic manipulation but often struggle with non-Markovian tasks requiring long-term memory and reasoning due to their reliance on immediate observations. Existing solutions face a frequency-competence paradox, where high-performance models are to…

Cited by 0SourceScholar
2026

LIBERO-Plus: A Progressive Robustness Benchmark for Visual-Language-Action Models

CVPR 2026

Visual-Language-Action (VLA) models report impressive success rates exceeding 95% on robotic manipulation benchmarks, yet these results may mask fundamental weaknesses in robustness. Current simulation-based robustness evaluations suffer from narrow perturbation coverage, manual design constraints,

Cited by 0SourcecodeScholar
2026

Learning to Move Before Learning to Do: Task-Agnostic pretraining for VLAs

ICML 2026poster

Vision-Language-Action (VLA) models are bottlenecked by the scarcity of expert demonstrations—expensive triplets of observations, language instructions, and actions. We propose that learning ''how to move'' can be decoupled from learning ''what to do,'' and that the former requires no task labels at…

Cited by 0SourceScholar
2026

SRPO: Self-Referential Policy Optimization for Vision-Language-Action Models

CVPR 2026

Vision-Language-Action (VLA) models excel in robotic manipulation but are constrained by their heavy reliance on expert demonstrations, leading to demonstration bias and limiting performance. Reinforcement learning (RL) is a vital post-training strategy to overcome these limits, yet current VLA-RL m

Cited by 0SourceScholar
2021

Evaluating Initialization Methods for Discriminative and Fast-Converging HGMM Point Clouds

ICRA 2021poster

Discriminative data representations for point cloud data are critical for computer vision applications. Recently, the Hierarchical Gaussian Mixture Model (HGMM) has become a popular representation due to its compactness and real-time execution. However, HGMM still lacks a well-designed and robust in…

Cited by 1SourceScholar