CoSTA: End-to-End Comprehensive Space-Time Entanglement for Spatio-Temporal Video Grounding
This paper studies the spatio-temporal video grounding task, which aims to localize a spatio-temporal tube in an untrimmed video based on the given text description of an event. Existing one-stage approaches suffer from insufficient space-time interaction in two aspects: i) less precise prediction o…