InRanker: Distilled Rankers for Zero-shot Information Retrieval

Laitz, Thiago; Papakostas, Konstantinos; Lotufo, Roberto; Nogueira, Rodrigo

Computer Science > Information Retrieval

arXiv:2401.06910 (cs)

[Submitted on 12 Jan 2024]

Title:InRanker: Distilled Rankers for Zero-shot Information Retrieval

Authors:Thiago Laitz, Konstantinos Papakostas, Roberto Lotufo, Rodrigo Nogueira

View PDF HTML (experimental)

Abstract:Despite multi-billion parameter neural rankers being common components of state-of-the-art information retrieval pipelines, they are rarely used in production due to the enormous amount of compute required for inference. In this work, we propose a new method for distilling large rankers into their smaller versions focusing on out-of-domain effectiveness. We introduce InRanker, a version of monoT5 distilled from monoT5-3B with increased effectiveness on out-of-domain scenarios. Our key insight is to use language models and rerankers to generate as much as possible synthetic "in-domain" training data, i.e., data that closely resembles the data that will be seen at retrieval time. The pipeline consists of two distillation phases that do not require additional user queries or manual annotations: (1) training on existing supervised soft teacher labels, and (2) training on teacher soft labels for synthetic queries generated using a large language model. Consequently, models like monoT5-60M and monoT5-220M improved their effectiveness by using the teacher's knowledge, despite being 50x and 13x smaller, respectively. Models and code are available at this https URL.

Subjects:	Information Retrieval (cs.IR)
Cite as:	arXiv:2401.06910 [cs.IR]
	(or arXiv:2401.06910v1 [cs.IR] for this version)
	https://doi.org/10.48550/arXiv.2401.06910

Submission history

From: Thiago Laitz [view email]
[v1] Fri, 12 Jan 2024 21:52:42 UTC (331 KB)

Computer Science > Information Retrieval

Title:InRanker: Distilled Rankers for Zero-shot Information Retrieval

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Information Retrieval

Title:InRanker: Distilled Rankers for Zero-shot Information Retrieval

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators