EMNLP 2021finding6 citations

Improving Embedding-based Large-scale Retrieval via Label Enhancement

Peiyang Liu, Xi Wang, Sen Wang, Wei Ye, Xiangyu Xi, Shikun Zhang

Abstract

Current embedding-based large-scale retrieval models are trained with 0-1 hard label that indicates whether a query is relevant to a document, ignoring rich information of the relevance degree. This paper proposes to improve embedding-based retrieval from the perspective of better characterizing the query-document relevance degree by introducing label enhancement (LE) for the first time. To generate label distribution in the retrieval scenario, we design a novel and effective supervised LE method that incorporates prior knowledge from dynamic term weighting methods into contextual embeddings. Our method significantly outperforms four competitive existing retrieval models and its counterparts equipped with two alternative LE techniques by training models with the generated label distribution as auxiliary supervision information. The superiority can be easily observed on English and Chinese large-scale retrieval tasks under both standard and cold-start settings.

BibTeX
@inproceedings{liu-etal-2021-improving-embedding-based,
    title = "Improving Embedding-based Large-scale Retrieval via Label Enhancement",
    author = "Liu, Peiyang  and
      Wang, Xi  and
      Wang, Sen  and
      Ye, Wei  and
      Xi, Xiangyu  and
      Zhang, Shikun",
    editor = "Moens, Marie-Francine  and
      Huang, Xuanjing  and
      Specia, Lucia  and
      Yih, Scott Wen-tau",
    booktitle = "Findings of the Association for Computational Linguistics: EMNLP 2021",
    month = nov,
    year = "2021",
    address = "Punta Cana, Dominican Republic",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2021.findings-emnlp.13/",
    doi = "10.18653/v1/2021.findings-emnlp.13",
    pages = "133--142"
}
Improving Embedding-based Large-scale Retrieval via Label Enhancement · EMNLP 2021