TMT-VIS: Taxonomy-aware Multi-dataset Joint Training for Video Instance Segmentation

Zheng, Rongkun; Qi, Lu; Chen, Xi; Wang, Yi; Wang, Kun; Qiao, Yu; Zhao, Hengshuang

Computer Science > Computer Vision and Pattern Recognition

arXiv:2312.06630 (cs)

[Submitted on 11 Dec 2023 (v1), last revised 17 Mar 2024 (this version, v3)]

Title:TMT-VIS: Taxonomy-aware Multi-dataset Joint Training for Video Instance Segmentation

Authors:Rongkun Zheng, Lu Qi, Xi Chen, Yi Wang, Kun Wang, Yu Qiao, Hengshuang Zhao

View PDF HTML (experimental)

Abstract:Training on large-scale datasets can boost the performance of video instance segmentation while the annotated datasets for VIS are hard to scale up due to the high labor cost. What we possess are numerous isolated filed-specific datasets, thus, it is appealing to jointly train models across the aggregation of datasets to enhance data volume and diversity. However, due to the heterogeneity in category space, as mask precision increases with the data volume, simply utilizing multiple datasets will dilute the attention of models on different taxonomies. Thus, increasing the data scale and enriching taxonomy space while improving classification precision is important. In this work, we analyze that providing extra taxonomy information can help models concentrate on specific taxonomy, and propose our model named Taxonomy-aware Multi-dataset Joint Training for Video Instance Segmentation (TMT-VIS) to address this vital challenge. Specifically, we design a two-stage taxonomy aggregation module that first compiles taxonomy information from input videos and then aggregates these taxonomy priors into instance queries before the transformer decoder. We conduct extensive experimental evaluations on four popular and challenging benchmarks, including YouTube-VIS 2019, YouTube-VIS 2021, OVIS, and UVO. Our model shows significant improvement over the baseline solutions, and sets new state-of-the-art records on all benchmarks. These appealing and encouraging results demonstrate the effectiveness and generality of our approach. The code is available at this https URL .

Comments:	NeurIPS 2023
Subjects:	Computer Vision and Pattern Recognition (cs.CV)
Cite as:	arXiv:2312.06630 [cs.CV]
	(or arXiv:2312.06630v3 [cs.CV] for this version)
	https://doi.org/10.48550/arXiv.2312.06630

Submission history

From: Rongkun Zheng [view email]
[v1] Mon, 11 Dec 2023 18:50:09 UTC (1,493 KB)
[v2] Tue, 12 Dec 2023 05:38:52 UTC (1,493 KB)
[v3] Sun, 17 Mar 2024 20:15:45 UTC (1,493 KB)

Computer Science > Computer Vision and Pattern Recognition

Title:TMT-VIS: Taxonomy-aware Multi-dataset Joint Training for Video Instance Segmentation

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computer Vision and Pattern Recognition

Title:TMT-VIS: Taxonomy-aware Multi-dataset Joint Training for Video Instance Segmentation

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators