Navigating the Unseen: Zero-shot Scene Graph Generation via Capsule-Based Equivariant Features
In scene graph generation (SGG), the accurate prediction of unseen triples is essential for its effectiveness in downstream vision-language tasks. We hypothesize that the predicates of unseen triples can be viewed as transformations of seen predicates in feature space, and the essence of the zero-sh…