Data augmentation methods for classifying Korean texts

Citations

WEB OF SCIENCE

0

초록

Data augmentation is widely adopted in computer vision. In contrast, research on data augmentation in thefield of natural language processing has been limited. We propose several data augmentation methods to supportthe classification of Korean texts. We increase the size and diversity of text data which are specifically tailoredto Korean. These methods adopt and adjust the existing data augmentation for English texts. We could improvethe classification accuracy and sometimes regularize the natural language models to reduce the overfits. Ourcontribution to the data augmentation regarding Korean texts compose of three parts. 1) data augmentation withSpelling Correction, 2) Easy data augmentation based on part-of-speech tagging, and 3) Data augmentation withconditional Masked Language Modeling. Our experiments show that classification accuracy can be improvedwith the aids of our proposed methods. Due to the limit of computing facilities, we consider rather small-scaleKorean texts only.

키워드

data augmentation; Korean text classification; masked language modeling; BERT
제목
Data augmentation methods for classifying Korean texts
저자
Jeon, Jihyun; Jung, Yoonsuh
DOI
10.5351/KJAS.2024.37.5.599
발행일
2024-10
유형
Article
저널명
응용통계연구
권
37
호
5