2026
From Scene to Object: Enhancing Open-Vocabulary Object Detection via Foreground-Background Context Reasoning
AAAI 2026technical
Open-Vocabulary Object Detection (OVOD) aims to detect both known and novel categories in complex visual scenes, surpassing the limitations of conventional closed-set detectors. Recent advances in vision-language models (VLMs) like CLIP have enabled zero-shot recognition by aligning visual features