ACL 2022findings15 citations

ED2LM: Encoder-Decoder to Language Model for Faster Document Re-ranking Inference

Kai Hui, Honglei Zhuang, Tao Chen, Zhen Qin, Jing Lu, Dara Bahri, Ji Ma, Jai Gupta

Abstract

State-of-the-art neural models typically encode document-query pairs using cross-attention for re-ranking. To this end, models generally utilize an encoder-only (like BERT) paradigm or an encoder-decoder (like T5) approach. These paradigms, however, are not without flaws, i.e., running the model on all query-document pairs at inference-time incurs a significant computational cost. This paper proposes a new training and inference paradigm for re-ranking. We propose to finetune a pretrained encoder-decoder model using in the form of document to query generation. Subsequently, we show that this encoder-decoder architecture can be decomposed into a decoder-only language model during inference. This results in significant inference time speedups since the decoder-only architecture only needs to learn to interpret static encoder embeddings during inference. Our experiments show that this new paradigm achieves results that are comparable to the more expensive cross-attention ranking approaches while being up to 6.8X faster. We believe this work paves the way for more efficient neural rankers that leverage large pretrained models.

BibTeX
@inproceedings{hui-etal-2022-ed2lm,
    title = "{ED}2{LM}: Encoder-Decoder to Language Model for Faster Document Re-ranking Inference",
    author = "Hui, Kai  and
      Zhuang, Honglei  and
      Chen, Tao  and
      Qin, Zhen  and
      Lu, Jing  and
      Bahri, Dara  and
      Ma, Ji  and
      Gupta, Jai  and
      Nogueira dos Santos, Cicero  and
      Tay, Yi  and
      Metzler, Donald",
    editor = "Muresan, Smaranda  and
      Nakov, Preslav  and
      Villavicencio, Aline",
    booktitle = "Findings of the Association for Computational Linguistics: ACL 2022",
    month = may,
    year = "2022",
    address = "Dublin, Ireland",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2022.findings-acl.295/",
    doi = "10.18653/v1/2022.findings-acl.295",
    pages = "3747--3758"
}
ED2LM: Encoder-Decoder to Language Model for Faster Document Re-ranking Inference · ACL 2022