COLING 2020main9 citations

SumTitles: a Summarization Dataset with Low Extractiveness

Valentin Malykh, Konstantin Chernis, Ekaterina Artemova, Irina Piontkovskaya

Abstract

The existing dialogue summarization corpora are significantly extractive. We introduce a methodology for dataset extractiveness evaluation and present a new low-extractive corpus of movie dialogues for abstractive text summarization along with baseline evaluation. The corpus contains 153k dialogues and consists of three parts: 1) automatically aligned subtitles, 2) automatically aligned scenes from scripts, and 3) manually aligned scenes from scripts. We also present an alignment algorithm which we use to construct the corpus.

BibTeX
@inproceedings{malykh-etal-2020-sumtitles,
    title = "{S}um{T}itles: a Summarization Dataset with Low Extractiveness",
    author = "Malykh, Valentin  and
      Chernis, Konstantin  and
      Artemova, Ekaterina  and
      Piontkovskaya, Irina",
    editor = "Scott, Donia  and
      Bel, Nuria  and
      Zong, Chengqing",
    booktitle = "Proceedings of the 28th International Conference on Computational Linguistics",
    month = dec,
    year = "2020",
    address = "Barcelona, Spain (Online)",
    publisher = "International Committee on Computational Linguistics",
    url = "https://aclanthology.org/2020.coling-main.503/",
    doi = "10.18653/v1/2020.coling-main.503",
    pages = "5718--5730"
}
SumTitles: a Summarization Dataset with Low Extractiveness · COLING 2020