Zero-shot voice conversion with HuBERT

Zero-shot voice conversion with HuBERT

초록

This study introduces an innovative model for zero-shot voice conversion that utilizes the capabilities of HuBERT. Zero-shot voice conversion models can transform the speech of one speaker to mimic that of another, even when the model has not been exposed to the target speaker's voice during the training phase. Comprising five main components (HuBERT, feature encoder, flow, speaker encoder, and vocoder), the model offers remarkable performance across a range of scenarios. Notably, it excels in the challenging unseen-to-unseen voice-conversion tasks. The effectiveness of the model was assessed based on the mean opinion scores and similarity scores, reflecting high voice quality and similarity to the target speakers. This model demonstrates considerable promise for a range of real-world applications demanding high-quality voice conversion. This study sets a precedent in the exploration of HuBERT-based models for voice conversion, and presents new directions for future research in this domain. Despite its complexities, the robust performance of this model underscores the viability of HuBERT in advancing voice conversion technology, making it a significant contributor to the field.

키워드

voice conversiondeep learningmachine learning
제목
Zero-shot voice conversion with HuBERT
제목 (타언어)
Zero-shot voice conversion with HuBERT
저자
정혜리남호성
발행일
2023-09
저널명
말소리와 음성과학
15
3
페이지
69 ~ 74