Recurrent Mixture Density Network for Spatiotemporal Visual Attention

Bazzani, Loris; Larochelle, Hugo; Torresani, Lorenzo

Computer Science > Computer Vision and Pattern Recognition

arXiv:1603.08199v2 (cs)

[Submitted on 27 Mar 2016 (v1), revised 3 Apr 2016 (this version, v2), latest version 11 Feb 2017 (v4)]

Title:Recurrent Mixture Density Network for Spatiotemporal Visual Attention

Authors:Loris Bazzani, Hugo Larochelle, Lorenzo Torresani

View PDF

Abstract:The high-dimensional and redundant nature of video have pushed researchers to seek the design of attentional models that can dynamically focus computations on the spatiotemporal volumes that are most relevant. Specifically, these models have been used to eliminate or down-weight background pixels that are not important for the task at hand. In order to deal with this problem, we propose an attentional model that learns where to look in a video directly from human fixation data. The proposed model leverages deep 3D convolutional features to represent clip segments in videos. This clip-level representation is aggregated over time by a long short-term memory network that connects into a mixture density network model of the likely positions of fixations in each frame. The resulting model is trained end to end using backpropagation. Our experiments show state-of-the-art performance on saliency prediction for videos. Experiments on Hollywood2 and UCF101 also show that the saliency can be used to improve classification accuracy on action recognition tasks.

Subjects:	Computer Vision and Pattern Recognition (cs.CV)
Cite as:	arXiv:1603.08199 [cs.CV]
	(or arXiv:1603.08199v2 [cs.CV] for this version)
	https://doi.org/10.48550/arXiv.1603.08199

Submission history

From: Loris Bazzani [view email]
[v1] Sun, 27 Mar 2016 10:34:22 UTC (609 KB)
[v2] Sun, 3 Apr 2016 14:17:51 UTC (620 KB)
[v3] Sun, 15 May 2016 11:55:35 UTC (626 KB)
[v4] Sat, 11 Feb 2017 10:05:06 UTC (776 KB)

Monday, May 5: arXiv will be READ ONLY at 9:00AM EST for approximately 30 minutes. We apologize for any inconvenience.

Computer Science > Computer Vision and Pattern Recognition

Title:Recurrent Mixture Density Network for Spatiotemporal Visual Attention

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computer Vision and Pattern Recognition

Title:Recurrent Mixture Density Network for Spatiotemporal Visual Attention

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators