AAAI 2024technical52 citations

Fine-Grained Distillation for Long Document Retrieval

Yucheng Zhou, Tao Shen, Xiubo Geng, Chongyang Tao, Jianbing Shen, Guodong Long, Can Xu, Daxin Jiang

Abstract

Long document retrieval aims to fetch query-relevant documents from a large-scale collection, where knowledge distillation has become de facto to improve a retriever by mimicking a heterogeneous yet powerful cross-encoder. However, in contrast to passages or sentences, retrieval on long documents suffers from the \textit{scope hypothesis} that a long document may cover multiple topics. This maximizes their structure heterogeneity and poses a granular-mismatch issue, leading to an inferior distillation efficacy. In this work, we propose a new learning framework, fine-grained distillation (FGD), for long-document retrievers. While preserving the conventional dense retrieval paradigm, it first produces global-consistent representations crossing different fine granularity and then applies multi-granular aligned distillation merely during training. In experiments, we evaluate our framework on two long-document retrieval benchmarks, which show state-of-the-art performance.

BibTeX
@article{Zhou_Shen_Geng_Tao_Shen_Long_Xu_Jiang_2024, title={Fine-Grained Distillation for Long Document Retrieval}, volume={38}, url={https://ojs.aaai.org/index.php/AAAI/article/view/29947}, DOI={10.1609/aaai.v38i17.29947}, abstractNote={Long document retrieval aims to fetch query-relevant documents from a large-scale collection, where knowledge distillation has become de facto to improve a retriever by mimicking a heterogeneous yet powerful cross-encoder. However, in contrast to passages or sentences, retrieval on long documents suffers from the \textit{scope hypothesis} that a long document may cover multiple topics. This maximizes their structure heterogeneity and poses a granular-mismatch issue, leading to an inferior distillation efficacy. In this work, we propose a new learning framework, fine-grained distillation (FGD), for long-document retrievers. While preserving the conventional dense retrieval paradigm, it first produces global-consistent representations crossing different fine granularity and then applies multi-granular aligned distillation merely during training. In experiments, we evaluate our framework on two long-document retrieval benchmarks, which show state-of-the-art performance.}, number={17}, journal={Proceedings of the AAAI Conference on Artificial Intelligence}, author={Zhou, Yucheng and Shen, Tao and Geng, Xiubo and Tao, Chongyang and Shen, Jianbing and Long, Guodong and Xu, Can and Jiang, Daxin}, year={2024}, month={Mar.}, pages={19732-19740} }
Fine-Grained Distillation for Long Document Retrieval · AAAI 2024