You Only Look at Once for Real-time and Generic Multi-Task

Wang, Jiayuan; Wu, Q. M. Jonathan; Zhang, Ning

Computer Science > Computer Vision and Pattern Recognition

arXiv:2310.01641 (cs)

[Submitted on 2 Oct 2023 (v1), last revised 24 Apr 2024 (this version, v4)]

Title:You Only Look at Once for Real-time and Generic Multi-Task

Authors:Jiayuan Wang, Q. M. Jonathan Wu, Ning Zhang

View PDF HTML (experimental)

Abstract:High precision, lightweight, and real-time responsiveness are three essential requirements for implementing autonomous driving. In this study, we incorporate A-YOLOM, an adaptive, real-time, and lightweight multi-task model designed to concurrently address object detection, drivable area segmentation, and lane line segmentation tasks. Specifically, we develop an end-to-end multi-task model with a unified and streamlined segmentation structure. We introduce a learnable parameter that adaptively concatenates features between necks and backbone in segmentation tasks, using the same loss function for all segmentation tasks. This eliminates the need for customizations and enhances the model's generalization capabilities. We also introduce a segmentation head composed only of a series of convolutional layers, which reduces the number of parameters and inference time. We achieve competitive results on the BDD100k dataset, particularly in visualization outcomes. The performance results show a mAP50 of 81.1% for object detection, a mIoU of 91.0% for drivable area segmentation, and an IoU of 28.8% for lane line segmentation. Additionally, we introduce real-world scenarios to evaluate our model's performance in a real scene, which significantly outperforms competitors. This demonstrates that our model not only exhibits competitive performance but is also more flexible and faster than existing multi-task models. The source codes and pre-trained models are released at this https URL

Subjects:	Computer Vision and Pattern Recognition (cs.CV)
Cite as:	arXiv:2310.01641 [cs.CV]
	(or arXiv:2310.01641v4 [cs.CV] for this version)
	https://doi.org/10.48550/arXiv.2310.01641

Submission history

From: Jiayuan Wang [view email]
[v1] Mon, 2 Oct 2023 21:09:43 UTC (9,334 KB)
[v2] Tue, 10 Oct 2023 03:50:28 UTC (9,444 KB)
[v3] Thu, 2 Nov 2023 16:52:42 UTC (9,446 KB)
[v4] Wed, 24 Apr 2024 20:05:04 UTC (8,776 KB)

Computer Science > Computer Vision and Pattern Recognition

Title:You Only Look at Once for Real-time and Generic Multi-Task

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computer Vision and Pattern Recognition

Title:You Only Look at Once for Real-time and Generic Multi-Task

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators