Many-to-many voice conversion experiments using a Korean speech corpus

Citations

SCOPUS

0

초록

Recently, Generative Adversarial Networks (GAN) and Variational AutoEncoders (VAE) have been applied to voice conversion that can make use of non-parallel training data. Especially, Conditional Cycle-Consistent Generative Adversarial Networks (CC-GAN) and Cycle-Consistent Variational AutoEncoders (CycleVAE) show promising results in many-to-many voice conversion among multiple speakers. However, the number of speakers has been relatively small in the conventional voice conversion studies using the CC-GANs and the CycleVAEs. In this paper, we extend the number of speakers to 100, and analyze the performances of the many-to-many voice conversion methods experimentally. It has been found through the experiments that the CC-GAN shows 4.5 % less Mel-Cepstral Distortion (MCD) for a small number of speakers, whereas the CycleVAE shows 12.7 % less MCD in a limited training time for a large number of speakers. Copyright © 2022 The Acoustical Society of Korea.

키워드

Conditional Cycle-Consistent Generative Adversarial Network (CC-GAN)Cycle-Consistent Variational AutoEncoder (CycleVAE)Generative Adversarial Network (GAN)Variational AutoEncoder (VAE)Voice conversion
제목
Many-to-many voice conversion experiments using a Korean speech corpus
저자
Yook, D.Seo, H.Ko, B.Yoo, I.-C.
DOI
10.7776/ASK.2022.41.3.351
발행일
2022
유형
Article
저널명
한국음향학회지
41
3
페이지
351 ~ 358