Similarity Search for Efficient Active Learning and Search of Rare Concepts

Coleman, Cody; Chou, Edward; Katz-Samuels, Julian; Culatana, Sean; Bailis, Peter; Berg, Alexander C.; Nowak, Robert; Sumbaly, Roshan; Zaharia, Matei; Yalniz, I. Zeki

Computer Science > Machine Learning

arXiv:2007.00077 (cs)

[Submitted on 30 Jun 2020 (v1), last revised 22 Jul 2021 (this version, v2)]

Title:Similarity Search for Efficient Active Learning and Search of Rare Concepts

Authors:Cody Coleman, Edward Chou, Julian Katz-Samuels, Sean Culatana, Peter Bailis, Alexander C. Berg, Robert Nowak, Roshan Sumbaly, Matei Zaharia, I. Zeki Yalniz

View PDF

Abstract:Many active learning and search approaches are intractable for large-scale industrial settings with billions of unlabeled examples. Existing approaches search globally for the optimal examples to label, scaling linearly or even quadratically with the unlabeled data. In this paper, we improve the computational efficiency of active learning and search methods by restricting the candidate pool for labeling to the nearest neighbors of the currently labeled set instead of scanning over all of the unlabeled data. We evaluate several selection strategies in this setting on three large-scale computer vision datasets: ImageNet, OpenImages, and a de-identified and aggregated dataset of 10 billion images provided by a large internet company. Our approach achieved similar mean average precision and recall as the traditional global approach while reducing the computational cost of selection by up to three orders of magnitude, thus enabling web-scale active learning.

Subjects:	Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (stat.ML)
Cite as:	arXiv:2007.00077 [cs.LG]
	(or arXiv:2007.00077v2 [cs.LG] for this version)
	https://doi.org/10.48550/arXiv.2007.00077

Submission history

From: Cody Coleman [view email]
[v1] Tue, 30 Jun 2020 19:46:10 UTC (1,263 KB)
[v2] Thu, 22 Jul 2021 16:54:12 UTC (1,879 KB)

Computer Science > Machine Learning

Title:Similarity Search for Efficient Active Learning and Search of Rare Concepts

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Machine Learning

Title:Similarity Search for Efficient Active Learning and Search of Rare Concepts

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators