Pre-train a Discriminative Text Encoder for Dense Retrieval via Contrastive Span Prediction

Ma, Xinyu; Guo, Jiafeng; Zhang, Ruqing; Fan, Yixing; Cheng, Xueqi

doi:10.1145/3477495.3531772

Computer Science > Information Retrieval

arXiv:2204.10641 (cs)

[Submitted on 22 Apr 2022]

Title:Pre-train a Discriminative Text Encoder for Dense Retrieval via Contrastive Span Prediction

Authors:Xinyu Ma, Jiafeng Guo, Ruqing Zhang, Yixing Fan, Xueqi Cheng

View PDF

Abstract:Dense retrieval has shown promising results in many information retrieval (IR) related tasks, whose foundation is high-quality text representation learning for effective search. Some recent studies have shown that autoencoder-based language models are able to boost the dense retrieval performance using a weak decoder. However, we argue that 1) it is not discriminative to decode all the input texts and, 2) even a weak decoder has the bypass effect on the encoder. Therefore, in this work, we introduce a novel contrastive span prediction task to pre-train the encoder alone, but still retain the bottleneck ability of the autoencoder. % Therefore, in this work, we propose to drop out the decoder and introduce a novel contrastive span prediction task to pre-train the encoder alone. The key idea is to force the encoder to generate the text representation close to its own random spans while far away from others using a group-wise contrastive loss. In this way, we can 1) learn discriminative text representations efficiently with the group-wise contrastive learning over spans and, 2) avoid the bypass effect of the decoder thoroughly. Comprehensive experiments over publicly available retrieval benchmark datasets show that our approach can outperform existing pre-training methods for dense retrieval significantly.

Comments:	Accepted to SIGIR 2022
Subjects:	Information Retrieval (cs.IR)
ACM classes:	H.3.3
Cite as:	arXiv:2204.10641 [cs.IR]
	(or arXiv:2204.10641v1 [cs.IR] for this version)
	https://doi.org/10.48550/arXiv.2204.10641
Related DOI:	https://doi.org/10.1145/3477495.3531772

Submission history

From: Xinyu Ma [view email]
[v1] Fri, 22 Apr 2022 11:22:29 UTC (1,221 KB)

Computer Science > Information Retrieval

Title:Pre-train a Discriminative Text Encoder for Dense Retrieval via Contrastive Span Prediction

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Information Retrieval

Title:Pre-train a Discriminative Text Encoder for Dense Retrieval via Contrastive Span Prediction

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators