Consistency-preserving Visual Question Answering in Medical Imaging

Tascon-Morales, Sergio; Márquez-Neila, Pablo; Sznitman, Raphael

Computer Science > Computer Vision and Pattern Recognition

arXiv:2206.13296 (cs)

[Submitted on 27 Jun 2022]

Title:Consistency-preserving Visual Question Answering in Medical Imaging

Authors:Sergio Tascon-Morales, Pablo Márquez-Neila, Raphael Sznitman

View PDF

Abstract:Visual Question Answering (VQA) models take an image and a natural-language question as input and infer the answer to the question. Recently, VQA systems in medical imaging have gained popularity thanks to potential advantages such as patient engagement and second opinions for clinicians. While most research efforts have been focused on improving architectures and overcoming data-related limitations, answer consistency has been overlooked even though it plays a critical role in establishing trustworthy models. In this work, we propose a novel loss function and corresponding training procedure that allows the inclusion of relations between questions into the training process. Specifically, we consider the case where implications between perception and reasoning questions are known a-priori. To show the benefits of our approach, we evaluate it on the clinically relevant task of Diabetic Macular Edema (DME) staging from fundus imaging. Our experiments show that our method outperforms state-of-the-art baselines, not only by improving model consistency, but also in terms of overall model accuracy. Our code and data are available at this https URL.

Comments:	Appears in Medical Image Computing and Computer Assisted Interventions (MICCAI), 2022
Subjects:	Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
Cite as:	arXiv:2206.13296 [cs.CV]
	(or arXiv:2206.13296v1 [cs.CV] for this version)
	https://doi.org/10.48550/arXiv.2206.13296

Submission history

From: Sergio Tascon Morales [view email]
[v1] Mon, 27 Jun 2022 13:38:50 UTC (6,942 KB)

Computer Science > Computer Vision and Pattern Recognition

Title:Consistency-preserving Visual Question Answering in Medical Imaging

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computer Vision and Pattern Recognition

Title:Consistency-preserving Visual Question Answering in Medical Imaging

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators