Hybrid-S2S: Video Object Segmentation with Recurrent Networks and Correspondence Matching

Azimi, Fatemeh; Frolov, Stanislav; Raue, Federico; Hees, Joern; Dengel, Andreas

Computer Science > Computer Vision and Pattern Recognition

arXiv:2010.05069 (cs)

[Submitted on 10 Oct 2020 (v1), last revised 7 Nov 2020 (this version, v2)]

Title:Hybrid-S2S: Video Object Segmentation with Recurrent Networks and Correspondence Matching

Authors:Fatemeh Azimi, Stanislav Frolov, Federico Raue, Joern Hees, Andreas Dengel

View PDF

Abstract:One-shot Video Object Segmentation~(VOS) is the task of pixel-wise tracking an object of interest within a video sequence, where the segmentation mask of the first frame is given at inference time. In recent years, Recurrent Neural Networks~(RNNs) have been widely used for VOS tasks, but they often suffer from limitations such as drift and error propagation. In this work, we study an RNN-based architecture and address some of these issues by proposing a hybrid sequence-to-sequence architecture named HS2S, utilizing a dual mask propagation strategy that allows incorporating the information obtained from correspondence matching. Our experiments show that augmenting the RNN with correspondence matching is a highly effective solution to reduce the drift problem. The additional information helps the model to predict more accurate masks and makes it robust against error propagation. We evaluate our HS2S model on the DAVIS2017 dataset as well as Youtube-VOS. On the latter, we achieve an improvement of 11.2pp in the overall segmentation accuracy over RNN-based state-of-the-art methods in VOS. We analyze our model's behavior in challenging cases such as occlusion and long sequences and show that our hybrid architecture significantly enhances the segmentation quality in these difficult scenarios.

Subjects:	Computer Vision and Pattern Recognition (cs.CV)
Cite as:	arXiv:2010.05069 [cs.CV]
	(or arXiv:2010.05069v2 [cs.CV] for this version)
	https://doi.org/10.48550/arXiv.2010.05069

Submission history

From: Fatemeh Azimi [view email]
[v1] Sat, 10 Oct 2020 19:00:43 UTC (37,295 KB)
[v2] Sat, 7 Nov 2020 09:33:51 UTC (37,305 KB)

Computer Science > Computer Vision and Pattern Recognition

Title:Hybrid-S2S: Video Object Segmentation with Recurrent Networks and Correspondence Matching

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computer Vision and Pattern Recognition

Title:Hybrid-S2S: Video Object Segmentation with Recurrent Networks and Correspondence Matching

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators