상세 보기
Many-to-many voice conversion experiments using a Korean speech corpus
- Yook, D.;
- Seo, H.;
- Ko, B.;
- Yoo, I.-C.
SCOPUS
0초록
Recently, Generative Adversarial Networks (GAN) and Variational AutoEncoders (VAE) have been applied to voice conversion that can make use of non-parallel training data. Especially, Conditional Cycle-Consistent Generative Adversarial Networks (CC-GAN) and Cycle-Consistent Variational AutoEncoders (CycleVAE) show promising results in many-to-many voice conversion among multiple speakers. However, the number of speakers has been relatively small in the conventional voice conversion studies using the CC-GANs and the CycleVAEs. In this paper, we extend the number of speakers to 100, and analyze the performances of the many-to-many voice conversion methods experimentally. It has been found through the experiments that the CC-GAN shows 4.5 % less Mel-Cepstral Distortion (MCD) for a small number of speakers, whereas the CycleVAE shows 12.7 % less MCD in a limited training time for a large number of speakers. Copyright © 2022 The Acoustical Society of Korea.
키워드
- 제목
- Many-to-many voice conversion experiments using a Korean speech corpus
- 저자
- Yook, D.; Seo, H.; Ko, B.; Yoo, I.-C.
- 발행일
- 2022
- 유형
- Article
- 저널명
- 한국음향학회지
- 권
- 41
- 호
- 3
- 페이지
- 351 ~ 358