TEINet: Towards an Efficient Architecture for Video Recognition

Liu, Zhaoyang; Luo, Donghao; Wang, Yabiao; Wang, Limin; Tai, Ying; Wang, Chengjie; Li, Jilin; Huang, Feiyue; Lu, Tong

Computer Science > Computer Vision and Pattern Recognition

arXiv:1911.09435 (cs)

[Submitted on 21 Nov 2019]

Title:TEINet: Towards an Efficient Architecture for Video Recognition

Authors:Zhaoyang Liu, Donghao Luo, Yabiao Wang, Limin Wang, Ying Tai, Chengjie Wang, Jilin Li, Feiyue Huang, Tong Lu

View PDF

Abstract:Efficiency is an important issue in designing video architectures for action recognition. 3D CNNs have witnessed remarkable progress in action recognition from videos. However, compared with their 2D counterparts, 3D convolutions often introduce a large amount of parameters and cause high computational cost. To relieve this problem, we propose an efficient temporal module, termed as Temporal Enhancement-and-Interaction (TEI Module), which could be plugged into the existing 2D CNNs (denoted by TEINet). The TEI module presents a different paradigm to learn temporal features by decoupling the modeling of channel correlation and temporal interaction. First, it contains a Motion Enhanced Module (MEM) which is to enhance the motion-related features while suppress irrelevant information (e.g., background). Then, it introduces a Temporal Interaction Module (TIM) which supplements the temporal contextual information in a channel-wise manner. This two-stage modeling scheme is not only able to capture temporal structure flexibly and effectively, but also efficient for model inference. We conduct extensive experiments to verify the effectiveness of TEINet on several benchmarks (e.g., Something-Something V1&V2, Kinetics, UCF101 and HMDB51). Our proposed TEINet can achieve a good recognition accuracy on these datasets but still preserve a high efficiency.

Comments:	Accepted by AAAI 2020
Subjects:	Computer Vision and Pattern Recognition (cs.CV)
Cite as:	arXiv:1911.09435 [cs.CV]
	(or arXiv:1911.09435v1 [cs.CV] for this version)
	https://doi.org/10.48550/arXiv.1911.09435

Submission history

From: Limin Wang [view email]
[v1] Thu, 21 Nov 2019 12:16:32 UTC (118 KB)

Computer Science > Computer Vision and Pattern Recognition

Title:TEINet: Towards an Efficient Architecture for Video Recognition

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computer Vision and Pattern Recognition

Title:TEINet: Towards an Efficient Architecture for Video Recognition

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators