Multi-View Attention Network for Visual Dialog

Park, Sungjin; Whang, Taesun; Yoon, Yeochan; Lim, Heuiseok

Computer Science > Artificial Intelligence

arXiv:2004.14025 (cs)

[Submitted on 29 Apr 2020 (v1), last revised 7 Oct 2020 (this version, v3)]

Title:Multi-View Attention Network for Visual Dialog

Authors:Sungjin Park, Taesun Whang, Yeochan Yoon, Heuiseok Lim

View PDF

Abstract:Visual dialog is a challenging vision-language task in which a series of questions visually grounded by a given image are answered. To resolve the visual dialog task, a high-level understanding of various multimodal inputs (e.g., question, dialog history, and image) is required. Specifically, it is necessary for an agent to 1) determine the semantic intent of question and 2) align question-relevant textual and visual contents among heterogeneous modality inputs. In this paper, we propose Multi-View Attention Network (MVAN), which leverages multiple views about heterogeneous inputs based on attention mechanisms. MVAN effectively captures the question-relevant information from the dialog history with two complementary modules (i.e., Topic Aggregation and Context Matching), and builds multimodal representations through sequential alignment processes (i.e., Modality Alignment). Experimental results on VisDial v1.0 dataset show the effectiveness of our proposed model, which outperforms the previous state-of-the-art methods with respect to all evaluation metrics.

Subjects:	Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
Cite as:	arXiv:2004.14025 [cs.AI]
	(or arXiv:2004.14025v3 [cs.AI] for this version)
	https://doi.org/10.48550/arXiv.2004.14025

Submission history

From: Sungjin Park [view email]
[v1] Wed, 29 Apr 2020 08:46:38 UTC (7,063 KB)
[v2] Tue, 6 Oct 2020 11:28:57 UTC (5,979 KB)
[v3] Wed, 7 Oct 2020 00:51:40 UTC (5,979 KB)

Computer Science > Artificial Intelligence

Title:Multi-View Attention Network for Visual Dialog

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Artificial Intelligence

Title:Multi-View Attention Network for Visual Dialog

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators