EfficientMFD: Towards More Efficient Multimodal Synchronous Fusion Detection

Zhang, Jiaqing; Cao, Mingxiang; Yang, Xue; Xie, Weiying; Lei, Jie; Li, Daixun; Yang, Geng; Huang, Wenbo; Li, Yunsong

Computer Science > Computer Vision and Pattern Recognition

arXiv:2403.09323v1 (cs)

[Submitted on 14 Mar 2024 (this version), latest version 23 May 2024 (v3)]

Title:EfficientMFD: Towards More Efficient Multimodal Synchronous Fusion Detection

Authors:Jiaqing Zhang, Mingxiang Cao, Xue Yang, Weiying Xie, Jie Lei, Daixun Li, Geng Yang, Wenbo Huang, Yunsong Li

View PDF HTML (experimental)

Abstract:Multimodal image fusion and object detection play a vital role in autonomous driving. Current joint learning methods have made significant progress in the multimodal fusion detection task combining the texture detail and objective semantic information. However, the tedious training steps have limited its applications to wider real-world industrial deployment. To address this limitation, we propose a novel end-to-end multimodal fusion detection algorithm, named EfficientMFD, to simplify models that exhibit decent performance with only one training step. Synchronous joint optimization is utilized in an end-to-end manner between two components, thus not being affected by the local optimal solution of the individual task. Besides, a comprehensive optimization is established in the gradient matrix between the shared parameters for both tasks. It can converge to an optimal point with fusion detection weights. We extensively test it on several public datasets, demonstrating superior performance on not only visually appealing fusion but also favorable detection performance (e.g., 6.6% mAP50:95) over other state-of-the-art approaches.

Subjects:	Computer Vision and Pattern Recognition (cs.CV)
Cite as:	arXiv:2403.09323 [cs.CV]
	(or arXiv:2403.09323v1 [cs.CV] for this version)
	https://doi.org/10.48550/arXiv.2403.09323

Submission history

From: Jiaqing Zhang [view email]
[v1] Thu, 14 Mar 2024 12:12:17 UTC (40,335 KB)
[v2] Tue, 21 May 2024 16:45:12 UTC (10,062 KB)
[v3] Thu, 23 May 2024 04:23:49 UTC (10,062 KB)

Computer Science > Computer Vision and Pattern Recognition

Title:EfficientMFD: Towards More Efficient Multimodal Synchronous Fusion Detection

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computer Vision and Pattern Recognition

Title:EfficientMFD: Towards More Efficient Multimodal Synchronous Fusion Detection

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators