AAAI 2026technical0 citations

VK-Det: Visual Knowledge Guided Prototype Learning for Open-Vocabulary Aerial Object Detection

Jianhang Yao, Yongbin Zheng, Siqi Lu, Wanying Xu, Peng Sun

Abstract

To identify objects beyond predefined categories, open-vocabulary aerial object detection (OVAD) leverages the zero-shot capabilities of visual-language models (VLMs) to generalize from base to novel categories. Existing approaches typically utilize self-learning mechanisms with weak text supervision to generate region-level pseudo-labels to align detectors with VLMs semantic spaces. However, text dependence induces semantic bias, restricting open-vocabulary expansion to text-specified concepts. We propose VK-Det, a visual knowledge-guided open-vocabulary object detection framework without extra supervision. First, we discover and leverage vision encoder

BibTeX
@inproceedings{aaai2026_vkdetvisualknowl,
  title = {VK-Det: Visual Knowledge Guided Prototype Learning for Open-Vocabulary Aerial Object Detection},
  author = {Jianhang Yao and Yongbin Zheng and Siqi Lu and Wanying Xu and Peng Sun},
  booktitle = {AAAI 2026},
  year = {2026}
}