Let's Split Up: Zero-Shot Classifier Edits for Fine-Grained Video Understanding
Video recognition models are typically trained on fixed taxonomies which are often too coarse, collapsing distinctions in object, manner or outcome under a single label. As tasks and definitions evolve, such models cannot accommodate emerging distinctions and collecting new annotations and retrainin…