HiLINK: Hierarchical linking of context-aware knowledge prediction and prompt tuning for bilingual knowledge-based visual question answering

Citations

WEB OF SCIENCE

1
Citations

SCOPUS

1

초록

Knowledge-based visual question answering (KBVQA) is a representative visual reasoning task that leverages external knowledge for question answering in situations where predicting the correct answer using only image and query data is difficult. In addition to KBVQA, various visual reasoning tasks have been actively studied for their potential to improve visual understanding by combining text and image modalities effectively. However, these tasks have primarily focused on high-resource languages, such as English. In contrast, studies on low-resource languages remain comparatively rare. To mitigate this research gap, we propose HiLINK, which utilizes multilingual data to enhance KBVQA performance in various languages. In this study, we use the BOK-VQA dataset to design the following key methodologies: We propose an end-to-end model that eliminates the need for a knowledge graph embedding-based training network by learning relationships between triplet knowledge components within prompts directly using Link-Tuning. We propose the HK-TriNet and HK-TriNet+ methodologies to perform triplet prediction based on contextualized knowledge relationships. Finally, we apply the frozen training approach as an alternative to conventional encoder joint training to improve the efficiency and performance of bilingual learning. HiLINK exhibits outstanding performance on the BOK-VQA dataset in three language configurations: bilingual, English, and Korean, outperforming the GEL-VQA method by +19.40%, +12.01%, and +11.30%, respectively. Furthermore, the effectiveness of the proposed method is validated based on a comprehensive analysis of bilingual embedding spaces, both visually and numerically. We expect this study to inspire future research on this topic and encourage practical applications of improved vision-language models.

키워드

Vision-language model; Knowledge-based visual question answering; Prompt learning; Multi-task learning
제목
HiLINK: Hierarchical linking of context-aware knowledge prediction and prompt tuning for bilingual knowledge-based visual question answering
저자
Jeong, Hyeonki; Kim, Taehyeong; Shin, Wooseok; Han, Sung Won
DOI
10.1016/j.knosys.2025.113556
발행일
2025-06-15
유형
Article
저널명
Knowledge-Based Systems
권
319