Where to Play: Retrieval of Video Segments using Natural-Language Queries

Lee, Sangkuk; Kim, Daesik; Lee, Myunggi; Hwang, Jihye; Kwak, Nojun

Computer Science > Computer Vision and Pattern Recognition

arXiv:1707.00251 (cs)

[Submitted on 2 Jul 2017]

Title:Where to Play: Retrieval of Video Segments using Natural-Language Queries

Authors:Sangkuk Lee, Daesik Kim, Myunggi Lee, Jihye Hwang, Nojun Kwak

View PDF

Abstract:In this paper, we propose a new approach for retrieval of video segments using natural language queries. Unlike most previous approaches such as concept-based methods or rule-based structured models, the proposed method uses image captioning model to construct sentential queries for visual information. In detail, our approach exploits multiple captions generated by visual features in each image with `Densecap'. Then, the similarities between captions of adjacent images are calculated, which is used to track semantically similar captions over multiple frames. Besides introducing this novel idea of 'tracking by captioning', the proposed method is one of the first approaches that uses a language generation model learned by neural networks to construct semantic query describing the relations and properties of visual information. To evaluate the effectiveness of our approach, we have created a new evaluation dataset, which contains about 348 segments of scenes in 20 movie-trailers. Through quantitative and qualitative evaluation, we show that our method is effective for retrieval of video segments using natural language queries.

Subjects:	Computer Vision and Pattern Recognition (cs.CV)
Cite as:	arXiv:1707.00251 [cs.CV]
	(or arXiv:1707.00251v1 [cs.CV] for this version)
	https://doi.org/10.48550/arXiv.1707.00251

Submission history

From: SangKuk Lee [view email]
[v1] Sun, 2 Jul 2017 07:56:06 UTC (1,834 KB)

Full-text links:

Access Paper:

view license

Current browse context:

cs.CV

< prev | next >

new | recent | 2017-07

Change to browse by:

References & Citations

DBLP - CS Bibliography

listing | bibtex

Sangkuk Lee
Daesik Kim
Myunggi Lee
Jihye Hwang
Nojun Kwak

export BibTeX citation

Computer Science > Computer Vision and Pattern Recognition

Title:Where to Play: Retrieval of Video Segments using Natural-Language Queries

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computer Vision and Pattern Recognition

Title:Where to Play: Retrieval of Video Segments using Natural-Language Queries

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators