3D CNNs with Adaptive Temporal Feature Resolutions

Fayyaz, Mohsen; Bahrami, Emad; Diba, Ali; Noroozi, Mehdi; Adeli, Ehsan; Van Gool, Luc; Gall, Juergen

Computer Science > Computer Vision and Pattern Recognition

arXiv:2011.08652 (cs)

[Submitted on 17 Nov 2020 (v1), last revised 11 Aug 2021 (this version, v4)]

Title:3D CNNs with Adaptive Temporal Feature Resolutions

Authors:Mohsen Fayyaz, Emad Bahrami, Ali Diba, Mehdi Noroozi, Ehsan Adeli, Luc Van Gool, Juergen Gall

View PDF

Abstract:While state-of-the-art 3D Convolutional Neural Networks (CNN) achieve very good results on action recognition datasets, they are computationally very expensive and require many GFLOPs. While the GFLOPs of a 3D CNN can be decreased by reducing the temporal feature resolution within the network, there is no setting that is optimal for all input clips. In this work, we therefore introduce a differentiable Similarity Guided Sampling (SGS) module, which can be plugged into any existing 3D CNN architecture. SGS empowers 3D CNNs by learning the similarity of temporal features and grouping similar features together. As a result, the temporal feature resolution is not anymore static but it varies for each input video clip. By integrating SGS as an additional layer within current 3D CNNs, we can convert them into much more efficient 3D CNNs with adaptive temporal feature resolutions (ATFR). Our evaluations show that the proposed module improves the state-of-the-art by reducing the computational cost (GFLOPs) by half while preserving or even improving the accuracy. We evaluate our module by adding it to multiple state-of-the-art 3D CNNs on various datasets such as Kinetics-600, Kinetics-400, mini-Kinetics, Something-Something V2, UCF101, and HMDB51.

Comments:	CVPR 2021
Subjects:	Computer Vision and Pattern Recognition (cs.CV)
Cite as:	arXiv:2011.08652 [cs.CV]
	(or arXiv:2011.08652v4 [cs.CV] for this version)
	https://doi.org/10.48550/arXiv.2011.08652

Submission history

From: Mohsen Fayyaz [view email]
[v1] Tue, 17 Nov 2020 14:34:05 UTC (1,112 KB)
[v2] Tue, 30 Mar 2021 13:06:10 UTC (2,908 KB)
[v3] Thu, 1 Apr 2021 10:31:57 UTC (2,933 KB)
[v4] Wed, 11 Aug 2021 09:14:20 UTC (2,953 KB)

Computer Science > Computer Vision and Pattern Recognition

Title:3D CNNs with Adaptive Temporal Feature Resolutions

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computer Vision and Pattern Recognition

Title:3D CNNs with Adaptive Temporal Feature Resolutions

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators