Orthogonal Gradient Penalty for Fast Training of Wasserstein GAN Based Multi-Task Autoencoder toward Robust Speech Recognition

  • Kao, Chao-Yuan
  • Park, Sangwook
  • Badi, Alzahra
  • Han, David K.
  • Ko, Hanseok
Citations

WEB OF SCIENCE

3
Citations

SCOPUS

3

초록

Performance in Automatic Speech Recognition (ASR) degrades dramatically in noisy environments. To alleviate this problem, a variety of deep networks based on convolutional neural networks and recurrent neural networks were proposed by applying L1 or L2 loss. In this Letter, we propose a new orthogonal gradient penalty (OGP) method for Wasserstein Generative Adversarial Networks (WGAN) applied to denoising and despeeching models. WGAN integrates a multi-task autoencoder which estimates not only speech features but also noise features from noisy speech. While achieving 14.1% improvement in Wasserstein distance convergence rate, the proposed OGP enhanced features are tested in ASR and achieve 9.7%, 8.6%, 6.2%, and 4.8% WER improvements over DDAE, MTAE, R-CED(CNN) and RNN models.

키워드

speech enhancementgenerative adversarial networksdeep learningrobust speech recognition
제목
Orthogonal Gradient Penalty for Fast Training of Wasserstein GAN Based Multi-Task Autoencoder toward Robust Speech Recognition
저자
Kao, Chao-YuanPark, SangwookBadi, AlzahraHan, David K.Ko, Hanseok
DOI
10.1587/transinf.2019EDL8183
발행일
2020-05
유형
Article
저널명
IEICE Transactions on Information and Systems
E103D
5
페이지
1195 ~ 1198