상세 보기
화자 임베딩과 발화 리듬의 연관관계에 대한 연구
- 김서현;
- 남호성
초록
The present study investigates if speech rhythm is encoded in the utterance-level speaker embedding which is an averaged value of frame-level speaker embeddings. When speaker encoders are used in Computer Assisted Pronunciation Training, finding what information is included in speaker embeddings is crucial because it defines what feature a learner should acquire to be fluent speaker. Rhythm has been regarded as a speaker identifiable feature. The speaker embeddings, however, may fail to capture rhythm features since the temporal dependency of prosody is likely to be lost by simple averaging. To quantify the degree to which rhythm information is encoded in the speaker embedding, the speaker embeddings were projected to the feature space by least square linear regression. The R-squared values for the rhythm features were consistently low across the models with the different number of parameters, in contrast to the acoustic features which showed the significantly high R-squared values. The result indicates that the utterance-mean embeddings did not encode speech rhythm of individual speaker. Based on the result, the way to better adopt speaker embeddings in CAPT system is discussed.
키워드
- 제목
- 화자 임베딩과 발화 리듬의 연관관계에 대한 연구
- 제목 (타언어)
- Does the Speaker Embedding Encode Speech Rhythm?
- 저자
- 김서현; 남호성
- 발행일
- 2021
- 저널명
- 외국어교육연구
- 권
- 35
- 호
- 2
- 페이지
- 131 ~ 144