상세 보기
초록
Language models (LMs) pretrained on a large text corpus and fine-tuned on a task data have a remarkable performance for document classification task. Recently, an adaptive pretraining method that re-pretrains the pretrained LMs using an additional dataset in the same domain with the given task to make up the domain discrepancy has reported significant performance improvements. However, current adaptive pretraining methods only focus on the domain gap between pretraining data and fine-tuning data. The writing style is also different because the pretraining data, e.g., Wikipedia, is written in a literary style, but the task data, e.g., customer review, is usually written in a colloquial style. In this work, we propose a colloquial-adaptive pretraining method that re-pretrains the pretrained LM with informal sentences to generalize the LM to colloquial style. We verify the proposed method based on multi-emotion classification datasets. The experimental results show that the proposed method yields improved classification performance on both low- and high-resource data.
키워드
- 제목
- 구어체 적응 사전 학습을 통한 한국어 감정 분류 성능 향상
- 제목 (타언어)
- Improving Korean Emotion Classification via Colloquial-Adaptive Pretraining
- 저자
- 이정훈; 김동화; 노영빈; 강필성
- 발행일
- 2021
- 저널명
- 대한산업공학회지
- 권
- 47
- 호
- 4
- 페이지
- 342 ~ 350