Learning to Diversify and Focus: A Reinforcement Framework for Open-Vocabulary HOI Detection
Open-Vocabulary Human-Object Interaction (OV-HOI) detection aims to recognize novel HOI categories beyond the training set. Existing OV-HOI detection approaches typically leverage CLIP to extract global visual representations and perform cross-attention between learnable queries and global features