상세 보기
Hierarchical Diffusion Model for Zero-Shot Singing Voice Synthesis With MIDI Priors
- Byun, Dong-Min;
- Kim, Seung-Bin;
- Lee, Seong-Whan
WEB OF SCIENCE
2초록
Singing voice synthesis systems have significantly advanced; however achieving high-quality singing voices in zero-shot tasks remains challenging. Traditional singing voice synthesis models face challenges in predicting the fundamental frequency (F0) of unseen speakers. In this study, we propose MIDI-Voice 2, which uses MIDI-driven priors to achieve high-quality singing voice synthesis, even in zero-shot tasks. We introduce a diffusion-based singing voice synthesis model that operates without F0. MIDI-Voice 2 consists of two diffusion models: a prior generator and a singing voice generator. The prior generator uses MIDI-driven priors, including accurate melody, to generate MIDI-style priors, and the singing voice generator uses these MIDI-style priors along with content and timbre information to generate singing voices. Disentangling the melody from timbre allows speaker adaptation without predicting the F0 of an unseen speaker. Additionally, we use a transformer-based diffusion to generate higher-quality audio in zero-shot tasks. Our experiments demonstrate that MIDI-Voice 2 improves speaker-adaptation without F0 and produces high-quality audio in zero-shot tasks.
키워드
- 제목
- Hierarchical Diffusion Model for Zero-Shot Singing Voice Synthesis With MIDI Priors
- 저자
- Byun, Dong-Min; Kim, Seung-Bin; Lee, Seong-Whan
- 발행일
- 2025
- 유형
- Article
- 권
- 33
- 페이지
- 2326 ~ 2336