Agent as Student: Learning From Informative Cues for Active Open-Vocabulary Recognition
Ruizhe Zeng, Lu Zhang, Xu Yang, Zhiyong Liu
Abstract
Active recognition, a fundamental task in embodied vision, aims to improve recognition performance by dynamically adjusting viewpoints and poses to mitigate the negative impacts of occlusion and blind spots. Although existing active recognition methods possess basic viewpoint adaptation capabilities, their performance is constrained by the insufficient exploitation of informative cues inherent in active recognition scenarios. To address this limitation, we propose an active open-vocabulary recognition method that enables agent to learn in a student-like manner by leveraging two types of informative cues. The first type involves failure experiences from previously encountered samples. We introduce a failure-driven learning mechanism that maintains a failure replay buffer and strategically replays failure cases based on the current competency of the agent. The second type is derived from high-quality frames in current tasks. We develop an adaptive feature fusion module that treats these frames as templates and incorporates an entropy-guided adaptation mechanism for online finetuning. Furthermore, we construct a benchmark dataset with diverse and fine-grained object categories to comprehensively evaluate active open-vocabulary recognition. Extensive experiments demonstrate superior performance of our method across multiple active open-vocabulary recognition benchmarks.
BibTeX
@inproceedings{ral2026_agentasstudentle,
title = {Agent as Student: Learning From Informative Cues for Active Open-Vocabulary Recognition},
author = {Ruizhe Zeng and Lu Zhang and Xu Yang and Zhiyong Liu},
booktitle = {RA-L 2026},
year = {2026}
}