EMNLP 2024main0 citations

GRIZAL: Generative Prior-guided Zero-Shot Temporal Action Localization

Onkar Kishor Susladkar, Gayatri Sudhir Deshmukh, Vandan Gorade, Sparsh Mittal

Abstract

Zero-shot temporal action localization (TAL) aims to temporally localize actions in videos without prior training examples. To address the challenges of TAL, we offer GRIZAL, a model that uses multimodal embeddings and dynamic motion cues to localize actions effectively. GRIZAL achieves sample diversity by using large-scale generative models such as GPT-4 for generating textual augmentations and DALL-E for generating image augmentations. Our model integrates vision-language embeddings with optical flow insights, optimized through a blend of supervised and self-supervised loss functions. On ActivityNet, Thumos14 and Charades-STA datasets, GRIZAL greatly outperforms state-of-the-art zero-shot TAL models, demonstrating its robustness and adaptability across a wide range of video content. We will make all the models and code publicly available by open-sourcing them.

BibTeX
@inproceedings{susladkar-etal-2024-grizal,
    title = "{GRIZAL}: Generative Prior-guided Zero-Shot Temporal Action Localization",
    author = "Susladkar, Onkar Kishor  and
      Deshmukh, Gayatri Sudhir  and
      Gorade, Vandan  and
      Mittal, Sparsh",
    editor = "Al-Onaizan, Yaser  and
      Bansal, Mohit  and
      Chen, Yun-Nung",
    booktitle = "Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing",
    month = nov,
    year = "2024",
    address = "Miami, Florida, USA",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2024.emnlp-main.1061/",
    doi = "10.18653/v1/2024.emnlp-main.1061",
    pages = "19046--19059"
}