상세 보기
Lip and Voice Synchronization Using Visual Attention
- 윤동련;
- 조현중
초록
This study explores lip-sync detection, focusing on the synchronization between lip movements and voices in videos. Typically, lip-syncdetection techniques involve cropping the facial area of a given video, utilizing the lower half of the cropped box as input for the visualencoder to extract visual features. To enhance the emphasis on the articulatory region of lips for more accurate lip-sync detection, wepropose utilizing a pre-trained visual attention-based encoder. The Visual Transformer Pooling (VTP) module is employed as the visualencoder, originally designed for the lip-reading task, predicting the script based solely on visual information without audio. Ourexperimental results demonstrate that, despite having fewer learning parameters, our proposed method outperforms the latest model,VocaList, on the LRS2 dataset, achieving a lip-sync detection accuracy of 94.5% based on five context frames. Moreover, our approachexhibits an approximately 8% superiority over VocaList in lip-sync detection accuracy, even on an untrained dataset, Acappella.
키워드
- 제목
- Lip and Voice Synchronization Using Visual Attention
- 제목 (타언어)
- Lip and Voice Synchronization Using Visual Attention
- 저자
- 윤동련; 조현중
- 발행일
- 2024-04
- 권
- 13
- 호
- 4
- 페이지
- 166 ~ 173