상세 보기
트랜스포머기반의 멀티모달 영상자막 생성요약
- 이민예;
- 한성원
초록
In this paper, we propose a MASTF methodology, which is a Multimodal Abstractive Summarization based on Transformer. Neural network models applied in the field of generative summaries utilizing conventional multi-modals were techniques utilizing hierarchical attention based on circulating neural networks. Although transformers showed excellent performance in various natural language processing fields, including generative summaries, there were no cases of application in multimodal-based generative summaries. Thus, in this paper, we use transformers to improve the performance of multimodal image subtitle generation summary models. Transformer-based models outperform hierarchical attention-based models by 24.17% on ROUGE-L basis and 10.52% on combining speech and text.
키워드
- 제목
- 트랜스포머기반의 멀티모달 영상자막 생성요약
- 제목 (타언어)
- Multi-Modal Abstractive Summarization based Transformer using Video Transcripts
- 저자
- 이민예; 한성원
- 발행일
- 2021
- 저널명
- 대한산업공학회지
- 권
- 47
- 호
- 5
- 페이지
- 433 ~ 443