Lucid Data Dreaming for Video Object Segmentation

Khoreva, Anna; Benenson, Rodrigo; Ilg, Eddy; Brox, Thomas; Schiele, Bernt

Computer Science > Computer Vision and Pattern Recognition

arXiv:1703.09554 (cs)

[Submitted on 28 Mar 2017 (v1), last revised 13 Mar 2019 (this version, v5)]

Title:Lucid Data Dreaming for Video Object Segmentation

Authors:Anna Khoreva, Rodrigo Benenson, Eddy Ilg, Thomas Brox, Bernt Schiele

View PDF

Abstract:Convolutional networks reach top quality in pixel-level video object segmentation but require a large amount of training data (1k~100k) to deliver such results. We propose a new training strategy which achieves state-of-the-art results across three evaluation datasets while using 20x~1000x less annotated data than competing methods. Our approach is suitable for both single and multiple object segmentation. Instead of using large training sets hoping to generalize across domains, we generate in-domain training data using the provided annotation on the first frame of each video to synthesize ("lucid dream") plausible future video frames. In-domain per-video training data allows us to train high quality appearance- and motion-based models, as well as tune the post-processing stage. This approach allows to reach competitive results even when training from only a single annotated frame, without ImageNet pre-training. Our results indicate that using a larger training set is not automatically better, and that for the video object segmentation task a smaller training set that is closer to the target domain is more effective. This changes the mindset regarding how many training samples and general "objectness" knowledge are required for the video object segmentation task.

Comments:	Accepted in International Journal of Computer Vision (IJCV)
Subjects:	Computer Vision and Pattern Recognition (cs.CV)
Cite as:	arXiv:1703.09554 [cs.CV]
	(or arXiv:1703.09554v5 [cs.CV] for this version)
	https://doi.org/10.48550/arXiv.1703.09554

Submission history

From: Anna Khoreva [view email]
[v1] Tue, 28 Mar 2017 12:56:40 UTC (6,705 KB)
[v2] Wed, 20 Sep 2017 12:50:09 UTC (13,201 KB)
[v3] Thu, 14 Dec 2017 13:38:20 UTC (9,382 KB)
[v4] Sun, 3 Feb 2019 16:34:23 UTC (9,100 KB)
[v5] Wed, 13 Mar 2019 19:55:04 UTC (8,982 KB)

Computer Science > Computer Vision and Pattern Recognition

Title:Lucid Data Dreaming for Video Object Segmentation

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computer Vision and Pattern Recognition

Title:Lucid Data Dreaming for Video Object Segmentation

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators