2023
NarrativeXL: a Large-scale Dataset for Long-Term Memory Models
EMNLP 2023long findings
We propose a new large-scale (nearly a million questions) ultra-long-context (more than 50,000 words average document length) reading comprehension dataset. Using GPT 3.5, we summarized each scene in 1,500 hand-curated fiction books from Project Gutenberg, which resulted in approximately 150 scene-l…