← Search

Yufeng Xu

2 accepted papers

2026

Good SFT Optimizes for SFT, Better SFT Prepares for Reinforcement Learning

ICML 2026poster

Post-training of reasoning LLMs is a holistic process that typically consists of an offline SFT stage followed by an online reinforcement learning (RL) stage. However, SFT is often optimized in isolation to maximize SFT performance alone. We show that, after identical RL training, models initialized…

Cited by 0SourceScholar
2024

BotanicGarden: A High-Quality Dataset for Robot Navigation in Unstructured Natural Environments

RA-L 2024

The rapid developments of mobile robotics and autonomous navigation over the years are largely empowered by public datasets for testing and upgrading, such as sensor odometry and SLAM tasks. Impressive demos and benchmark scores have arisen, which may suggest the maturity of existing navigation tech

Cited by 60SourcecodeScholar