EMNLP 2021finding20 citations

SciXGen: A Scientific Paper Dataset for Context-Aware Text Generation

Hong Chen, Hiroya Takamura, Hideki Nakayama

Abstract

Generating texts in scientific papers requires not only capturing the content contained within the given input but also frequently acquiring the external information called context. We push forward the scientific text generation by proposing a new task, namely context-aware text generation in the scientific domain, aiming at exploiting the contributions of context in generated texts. To this end, we present a novel challenging large-scale Scientific Paper Dataset for ConteXt-Aware Text Generation (SciXGen), consisting of well-annotated 205,304 papers with full references to widely-used objects (e.g., tables, figures, algorithms) in a paper. We comprehensively benchmark, using state-of-the-arts, the efficacy of our newly constructed SciXGen dataset in generating description and paragraph. Our dataset and benchmarks will be made publicly available to hopefully facilitate the scientific text generation research.

BibTeX
@inproceedings{chen-etal-2021-scixgen-scientific,
    title = "{S}ci{XG}en: A Scientific Paper Dataset for Context-Aware Text Generation",
    author = "Chen, Hong  and
      Takamura, Hiroya  and
      Nakayama, Hideki",
    editor = "Moens, Marie-Francine  and
      Huang, Xuanjing  and
      Specia, Lucia  and
      Yih, Scott Wen-tau",
    booktitle = "Findings of the Association for Computational Linguistics: EMNLP 2021",
    month = nov,
    year = "2021",
    address = "Punta Cana, Dominican Republic",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2021.findings-emnlp.128/",
    doi = "10.18653/v1/2021.findings-emnlp.128",
    pages = "1483--1492"
}
SciXGen: A Scientific Paper Dataset for Context-Aware Text Generation · EMNLP 2021