NAACL 2025industry0 citations

QSpell 250K: A Large-Scale, Practical Dataset for Chinese Search Query Spell Correction

Dezhi Ye, Haomei Jia, Junwei Hu, Tian Bowen, Jie Liu, Haijin Liang, Jin Ma, Wenmin Wang

Abstract

Chinese Search Query Spell Correction is a task designed to autonomously identify and correct typographical errors within queries in the search engine. Despite the availability of comprehensive datasets like Microsoft Speller and Webis, their monolingual nature and limited scope pose significant challenges in evaluating modern pre-trained language models such as BERT and GPT. To address this, we introduce QSpell 250K, a large-scale benchmark specifically developed for Chinese Query Spelling Correction. QSpell 250K offers several advantages: 1) It contains over 250K samples, which is ten times more than previous datasets. 2) It covers a broad range of topics, from formal entities to everyday colloquialisms and idiomatic expressions. 3) It includes both Chinese and English, addressing the complexities of code-switching. Each query undergoes three rounds of high-fidelity annotation to ensure accuracy. Our extensive testing across three popular models demonstrates that QSpell 250K effectively evaluates the efficacy of representative spelling correctors. We believe that QSpell 250K will significantly advance spelling correction methodologies. The accompanying data and code will be made publicly available.

BibTeX
@inproceedings{ye-etal-2025-qspell,
    title = "{QS}pell 250{K}: A Large-Scale, Practical Dataset for {C}hinese Search Query Spell Correction",
    author = "Ye, Dezhi  and
      Jia, Haomei  and
      Hu, Junwei  and
      Bowen, Tian  and
      Liu, Jie  and
      Liang, Haijin  and
      Ma, Jin  and
      Wang, Wenmin",
    editor = "Chen, Weizhu  and
      Yang, Yi  and
      Kachuee, Mohammad  and
      Fu, Xue-Yong",
    booktitle = "Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 3: Industry Track)",
    month = apr,
    year = "2025",
    address = "Albuquerque, New Mexico",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2025.naacl-industry.13/",
    pages = "148--155",
    ISBN = "979-8-89176-194-0"
}
QSpell 250K: A Large-Scale, Practical Dataset for Chinese Search Query Spell Correction · NAACL 2025