← Search

Yuanzhi Liu

2 accepted papers

2026

MMBench-Live: A Continuously Evolving Benchmark for Multimodal Models

ICML 2026poster

Evaluation benchmarks play a central role in assessing vision–language models (VLMs). However, most existing multimodal benchmarks are static, making them increasingly vulnerable to data contamination, temporal staleness, and high construction costs. In this work, we introduce MMBench-Live, a multi-…

Cited by 0SourceScholar
2024

BotanicGarden: A High-Quality Dataset for Robot Navigation in Unstructured Natural Environments

RA-L 2024

The rapid developments of mobile robotics and autonomous navigation over the years are largely empowered by public datasets for testing and upgrading, such as sensor odometry and SLAM tasks. Impressive demos and benchmark scores have arisen, which may suggest the maturity of existing navigation tech

Cited by 60SourcecodeScholar