구어체 적응 사전 학습을 통한 한국어 감정 분류 성능 향상

Improving Korean Emotion Classification via Colloquial-Adaptive Pretraining
  • 이정훈
  • 김동화
  • 노영빈
  • 강필성

초록

Language models (LMs) pretrained on a large text corpus and fine-tuned on a task data have a remarkable performance for document classification task. Recently, an adaptive pretraining method that re-pretrains the pretrained LMs using an additional dataset in the same domain with the given task to make up the domain discrepancy has reported significant performance improvements. However, current adaptive pretraining methods only focus on the domain gap between pretraining data and fine-tuning data. The writing style is also different because the pretraining data, e.g., Wikipedia, is written in a literary style, but the task data, e.g., customer review, is usually written in a colloquial style. In this work, we propose a colloquial-adaptive pretraining method that re-pretrains the pretrained LM with informal sentences to generalize the LM to colloquial style. We verify the proposed method based on multi-emotion classification datasets. The experimental results show that the proposed method yields improved classification performance on both low- and high-resource data.

키워드

Natural Language ProcessingTransfer LearningAdaptive PretrainingMulti-Emotion Classification
제목
구어체 적응 사전 학습을 통한 한국어 감정 분류 성능 향상
제목 (타언어)
Improving Korean Emotion Classification via Colloquial-Adaptive Pretraining
저자
이정훈김동화노영빈강필성
발행일
2021
저널명
대한산업공학회지
47
4
페이지
342 ~ 350