A Simple Cache Model for Image Recognition

Orhan, A. Emin

Computer Science > Computer Vision and Pattern Recognition

arXiv:1805.08709 (cs)

[Submitted on 21 May 2018 (v1), last revised 26 Nov 2018 (this version, v2)]

Title:A Simple Cache Model for Image Recognition

Authors:A. Emin Orhan

View PDF

Abstract:Training large-scale image recognition models is computationally expensive. This raises the question of whether there might be simple ways to improve the test performance of an already trained model without having to re-train or fine-tune it with new data. Here, we show that, surprisingly, this is indeed possible. The key observation we make is that the layers of a deep network close to the output layer contain independent, easily extractable class-relevant information that is not contained in the output layer itself. We propose to extract this extra class-relevant information using a simple key-value cache memory to improve the classification performance of the model at test time. Our cache memory is directly inspired by a similar cache model previously proposed for language modeling (Grave et al., 2017). This cache component does not require any training or fine-tuning; it can be applied to any pre-trained model and, by properly setting only two hyper-parameters, leads to significant improvements in its classification performance. Improvements are observed across several architectures and datasets. In the cache component, using features extracted from layers close to the output (but not from the output layer itself) as keys leads to the largest improvements. Concatenating features from multiple layers to form keys can further improve performance over using single-layer features as keys. The cache component also has a regularizing effect, a simple consequence of which is that it substantially increases the robustness of models against adversarial attacks.

Comments:	Published as a conference paper at NIPS 2018
Subjects:	Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG); Neural and Evolutionary Computing (cs.NE); Machine Learning (stat.ML)
Cite as:	arXiv:1805.08709 [cs.CV]
	(or arXiv:1805.08709v2 [cs.CV] for this version)
	https://doi.org/10.48550/arXiv.1805.08709

Submission history

From: Emin Orhan [view email]
[v1] Mon, 21 May 2018 17:50:14 UTC (481 KB)
[v2] Mon, 26 Nov 2018 18:24:18 UTC (516 KB)

Computer Science > Computer Vision and Pattern Recognition

Title:A Simple Cache Model for Image Recognition

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computer Vision and Pattern Recognition

Title:A Simple Cache Model for Image Recognition

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators