Exploring Triple Knowledge Cues for Zero-Shot Human-Object Interaction Detection
Current zero-shot human-object interaction detection methods often follow a two-phase pipeline, which uses a pre-trained detector to detect instances and then adopts CLIP to perform interaction prediction. During the second phase, they either obtain pairwise representations by directly performing Ro…