Image classification and captioning model considering a CAM-based disagreement loss

  • Yoon, Yeo Chan
  • Park, So Young
  • Park, Soo Myoung
  • Lim, Heuiseok
Citations

WEB OF SCIENCE

6
Citations

SCOPUS

8

초록

Image captioning has received significant interest in recent years, and notable results have been achieved. Most previous approaches have focused on generating visual descriptions from images, whereas a few approaches have exploited visual descriptions for image classification. This study demonstrates that a good performance can be achieved for both description generation and image classification through an end-to-end joint learning approach with a loss function, which encourages each task to reach a consensus. When given images and visual descriptions, the proposed model learns a multimodal intermediate embedding, which can represent both the textual and visual characteristics of an object. The performance can be improved for both tasks by sharing the multimodal embedding. Through a novel loss function based on class activation mapping, which localizes the discriminative image region of a model, we achieve a higher score when the captioning and classification model reaches a consensus on the key parts of the object. Using the proposed model, we established a substantially improved performance for each task on the UCSD Birds and Oxford Flowers datasets.

키워드

deep learningimage captioningimage classification
제목
Image classification and captioning model considering a CAM-based disagreement loss
저자
Yoon, Yeo ChanPark, So YoungPark, Soo MyoungLim, Heuiseok
DOI
10.4218/etrij.2018-0621
발행일
2020-02
유형
Article
저널명
ETRI Journal
42
1
페이지
67 ~ 77