Lip and Voice Synchronization Using Visual Attention

Lip and Voice Synchronization Using Visual Attention

초록

This study explores lip-sync detection, focusing on the synchronization between lip movements and voices in videos. Typically, lip-syncdetection techniques involve cropping the facial area of a given video, utilizing the lower half of the cropped box as input for the visualencoder to extract visual features. To enhance the emphasis on the articulatory region of lips for more accurate lip-sync detection, wepropose utilizing a pre-trained visual attention-based encoder. The Visual Transformer Pooling (VTP) module is employed as the visualencoder, originally designed for the lip-reading task, predicting the script based solely on visual information without audio. Ourexperimental results demonstrate that, despite having fewer learning parameters, our proposed method outperforms the latest model,VocaList, on the LRS2 dataset, achieving a lip-sync detection accuracy of 94.5% based on five context frames. Moreover, our approachexhibits an approximately 8% superiority over VocaList in lip-sync detection accuracy, even on an untrained dataset, Acappella.

키워드

입술-음성 동기화시각적 어텐션트랜스포머Lip-Voice SynchronizationVisual AttentionMulti-Modal Transformer
제목
Lip and Voice Synchronization Using Visual Attention
제목 (타언어)
Lip and Voice Synchronization Using Visual Attention
저자
윤동련조현중
발행일
2024-04
저널명
정보처리학회논문지. 소프트웨어 및 데이터 공학
13
4
페이지
166 ~ 173