COLING 2024main5 citations

Frame2: A FrameNet-based Multimodal Dataset for Tackling Text-image Interactions in Video

Frederico Belcavello, Tiago Timponi Torrent, Ely E. Matos, Adriana S. Pagano, Maucha Gamonal, Natalia Sigiliano, Lívia Vicente Dutra, Helen de Andrade Abreu

Abstract

This paper presents the Frame2 dataset, a multimodal dataset built from a corpus of a Brazilian travel TV show annotated for FrameNet categories for both the text and image communicative modes. Frame2 comprises 230 minutes of video, which are correlated with 2,915 sentences either transcribing the audio spoken during the episodes or the subtitling segments of the show where the host conducts interviews in English. For this first release of the dataset, a total of 11,796 annotation sets for the sentences and 6,841 for the video are included. Each of the former includes a target lexical unit evoking a frame or one or more frame elements. For each video annotation, a bounding box in the image is correlated with a frame, a frame element and lexical unit evoking a frame in FrameNet.

BibTeX
@inproceedings{belcavello-etal-2024-frame2,
    title = "Frame2: A {F}rame{N}et-based Multimodal Dataset for Tackling Text-image Interactions in Video",
    author = "Belcavello, Frederico  and
      Timponi Torrent, Tiago  and
      Matos, Ely E.  and
      Pagano, Adriana S.  and
      Gamonal, Maucha  and
      Sigiliano, Natalia  and
      Dutra, L{\'i}via Vicente  and
      de Andrade Abreu, Helen  and
      Samagaio, Mairon  and
      Carvalho, Mariane  and
      Campos, Franciany  and
      Azalim, Gabrielly  and
      Mazzei, Bruna  and
      de Oliveira, Mateus Fonseca  and
      Lo{\c{c}}asso Luz, Ana Carolina  and
      P{\'a}dua Ruiz, L{\'i}via  and
      Bellei, J{\'u}lia  and
      Pestana, Amanda  and
      Costa, Josiane  and
      Rabelo, Iasmin  and
      Silva, Anna Beatriz  and
      Roza, Raquel  and
      Souza, Mariana  and
      Oliveira, Igor",
    editor = "Calzolari, Nicoletta  and
      Kan, Min-Yen  and
      Hoste, Veronique  and
      Lenci, Alessandro  and
      Sakti, Sakriani  and
      Xue, Nianwen",
    booktitle = "Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024)",
    month = may,
    year = "2024",
    address = "Torino, Italia",
    publisher = "ELRA and ICCL",
    url = "https://aclanthology.org/2024.lrec-main.655/",
    pages = "7429--7437"
}
Frame2: A FrameNet-based Multimodal Dataset for Tackling Text-image Interactions in Video · COLING 2024