Prophet Attention: Predicting Attention with Future Attention for Image Captioning

Liu, Fenglin; Ren, Xuancheng; Wu, Xian; Fan, Wei; Zou, Yuexian; Sun, Xu

Computer Science > Computer Vision and Pattern Recognition

arXiv:2210.10914 (cs)

[Submitted on 19 Oct 2022 (v1), last revised 11 Apr 2023 (this version, v2)]

Title:Prophet Attention: Predicting Attention with Future Attention for Image Captioning

Authors:Fenglin Liu, Xuancheng Ren, Xian Wu, Wei Fan, Yuexian Zou, Xu Sun

View PDF

Abstract:Recently, attention based models have been used extensively in many sequence-to-sequence learning systems. Especially for image captioning, the attention based models are expected to ground correct image regions with proper generated words. However, for each time step in the decoding process, the attention based models usually use the hidden state of the current input to attend to the image regions. Under this setting, these attention models have a "deviated focus" problem that they calculate the attention weights based on previous words instead of the one to be generated, impairing the performance of both grounding and captioning. In this paper, we propose the Prophet Attention, similar to the form of self-supervision. In the training stage, this module utilizes the future information to calculate the "ideal" attention weights towards image regions. These calculated "ideal" weights are further used to regularize the "deviated" attention. In this manner, image regions are grounded with the correct words. The proposed Prophet Attention can be easily incorporated into existing image captioning models to improve their performance of both grounding and captioning. The experiments on the Flickr30k Entities and the MSCOCO datasets show that the proposed Prophet Attention consistently outperforms baselines in both automatic metrics and human evaluations. It is worth noticing that we set new state-of-the-arts on the two benchmark datasets and achieve the 1st place on the leaderboard of the online MSCOCO benchmark in terms of the default ranking score, i.e., CIDEr-c40.

Comments:	Accepted by NeurIPS 2020
Subjects:	Computer Vision and Pattern Recognition (cs.CV); Computation and Language (cs.CL)
Cite as:	arXiv:2210.10914 [cs.CV]
	(or arXiv:2210.10914v2 [cs.CV] for this version)
	https://doi.org/10.48550/arXiv.2210.10914

Submission history

From: Fenglin Liu [view email]
[v1] Wed, 19 Oct 2022 22:29:31 UTC (2,434 KB)
[v2] Tue, 11 Apr 2023 06:12:01 UTC (2,434 KB)

Computer Science > Computer Vision and Pattern Recognition

Title:Prophet Attention: Predicting Attention with Future Attention for Image Captioning

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computer Vision and Pattern Recognition

Title:Prophet Attention: Predicting Attention with Future Attention for Image Captioning

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators