A Large-Scale Chinese Long-Text Extractive Summarization Corpus
Kai Chen, Guanyu Fu, Qingcai Chen, Baotian Hu
Abstract
Recently, large-scale datasets have vastly facilitated the development in nearly domains of Natural Language Processing. However, lacking large scale Chinese corpus is still a critical bottleneck for further research on deep text summarization methods. In this paper, we publish a large-scale Chinese Long-text Extractive Summarization corpus named CLES. The CLES contains about 104K <summary, article> pairs, which is originally collected from Sina Weibo <sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">1</sup> . To verify the quality of the corpus, we also manually tagged the relevance score of 5,000 <summary, article> pairs. Our benchmark models on the proposed corpus include conventional deep learning based extractive models and several pre-trained Bert-based algorithms. Their performances are reported and briefly analyzed to facilitate further research on the corpus. We will release this corpus for further research <sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">2</sup> .
BibTeX
@inproceedings{icassp2021_alargescalechine,
title = {A Large-Scale Chinese Long-Text Extractive Summarization Corpus},
author = {Kai Chen and Guanyu Fu and Qingcai Chen and Baotian Hu},
booktitle = {ICASSP 2021},
year = {2021}
}