Meaning guided video captioning

Babariya, Rushi J.; Tamaki, Toru

Computer Science > Computer Vision and Pattern Recognition

arXiv:1912.05730 (cs)

[Submitted on 12 Dec 2019]

Title:Meaning guided video captioning

Authors:Rushi J. Babariya, Toru Tamaki

View PDF

Abstract:Current video captioning approaches often suffer from problems of missing objects in the video to be described, while generating captions semantically similar with ground truth sentences. In this paper, we propose a new approach to video captioning that can describe objects detected by object detection, and generate captions having similar meaning with correct captions. Our model relies on S2VT, a sequence-to-sequence model for video captioning. Given a sequence of video frames, the encoding RNN takes a frame as well as detected objects in the frame in order to incorporate the information of the objects in the scene. The following decoding RNN outputs are then fed into an attention layer and then to a decoder for generating captions. The caption is compared with the ground truth by learning metric so that vector representations of generated captions are semantically similar to those of ground truth. Experimental results with the MSDV dataset demonstrate that the performance of the proposed approach is much better than the model without the proposed meaning-guided framework, showing the effectiveness of the proposed model. Code are publicly available at this https URL.

Comments:	The 5th Asian Conference on Pattern Recognition (ACPR 2019)
Subjects:	Computer Vision and Pattern Recognition (cs.CV)
Cite as:	arXiv:1912.05730 [cs.CV]
	(or arXiv:1912.05730v1 [cs.CV] for this version)
	https://doi.org/10.48550/arXiv.1912.05730

Submission history

From: Toru Tamaki [view email]
[v1] Thu, 12 Dec 2019 02:05:45 UTC (118 KB)

Computer Science > Computer Vision and Pattern Recognition

Title:Meaning guided video captioning

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computer Vision and Pattern Recognition

Title:Meaning guided video captioning

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators