Detailed Information

Cited 0 time in webofscience Cited 0 time in scopus
Metadata Downloads

Multi-View Attention Network for Visual Dialog

Full metadata record
DC Field Value Language
dc.contributor.authorPark, Sungjin-
dc.contributor.authorWhang, Taesun-
dc.contributor.authorYoon, Yeochan-
dc.contributor.authorLim, Heuiseok-
dc.date.accessioned2021-11-22T08:40:21Z-
dc.date.available2021-11-22T08:40:21Z-
dc.date.created2021-08-30-
dc.date.issued2021-04-
dc.identifier.issn2076-3417-
dc.identifier.urihttps://scholar.korea.ac.kr/handle/2021.sw.korea/128333-
dc.description.abstractVisual dialog is a challenging vision-language task in which a series of questions visually grounded by a given image are answered. To resolve the visual dialog task, a high-level understanding of various multimodal inputs (e.g., question, dialog history, and image) is required. Specifically, it is necessary for an agent to (1) determine the semantic intent of question and (2) align question-relevant textual and visual contents among heterogeneous modality inputs. In this paper, we propose Multi-View Attention Network (MVAN), which leverages multiple views about heterogeneous inputs based on attention mechanisms. MVAN effectively captures the question-relevant information from the dialog history with two complementary modules (i.e., Topic Aggregation and Context Matching), and builds multimodal representations through sequential alignment processes (i.e., Modality Alignment). Experimental results on VisDial v1.0 dataset show the effectiveness of our proposed model, which outperforms previous state-of-the-art methods under both single model and ensemble settings.-
dc.languageEnglish-
dc.language.isoen-
dc.publisherMDPI-
dc.titleMulti-View Attention Network for Visual Dialog-
dc.typeArticle-
dc.contributor.affiliatedAuthorLim, Heuiseok-
dc.identifier.doi10.3390/app11073009-
dc.identifier.scopusid2-s2.0-85103855133-
dc.identifier.wosid000638347400001-
dc.identifier.bibliographicCitationAPPLIED SCIENCES-BASEL, v.11, no.7-
dc.relation.isPartOfAPPLIED SCIENCES-BASEL-
dc.citation.titleAPPLIED SCIENCES-BASEL-
dc.citation.volume11-
dc.citation.number7-
dc.type.rimsART-
dc.type.docTypeArticle-
dc.description.journalClass1-
dc.description.journalRegisteredClassscie-
dc.description.journalRegisteredClassscopus-
dc.relation.journalResearchAreaChemistry-
dc.relation.journalResearchAreaEngineering-
dc.relation.journalResearchAreaMaterials Science-
dc.relation.journalResearchAreaPhysics-
dc.relation.journalWebOfScienceCategoryChemistry, Multidisciplinary-
dc.relation.journalWebOfScienceCategoryEngineering, Multidisciplinary-
dc.relation.journalWebOfScienceCategoryMaterials Science, Multidisciplinary-
dc.relation.journalWebOfScienceCategoryPhysics, Applied-
dc.subject.keywordAuthorvisual dialog-
dc.subject.keywordAuthorattention mechanism-
dc.subject.keywordAuthormultimodal learning-
dc.subject.keywordAuthorvision-language-
Files in This Item
There are no files associated with this item.
Appears in
Collections
Graduate School > Department of Computer Science and Engineering > 1. Journal Articles

qrcode

Items in ScholarWorks are protected by copyright, with all rights reserved, unless otherwise indicated.

Altmetrics

Total Views & Downloads

BROWSE