최소대립 문장쌍을 활용한 한국어 사전학습모델의 통사 연구 활용 가능성 검증

Verification of Korean Pre-trained Models' Feasibility of Syntactic Research Using Pairwise Sentences

초록

Syntactic studies make use of the minimally pairwise sentences as an argumentation tool, because the pairs allow us to pay attention to the constraints of interest. Likewise, it is helpful to use a set of minimal pairs in deep learning-based experiments for assessing the syntactic ability of neural language models. In this context, this study verifies whether the deep learning Korean model has the ability to properly distinguish the well-formed expressions and the corresponding ill-formed expressions. In the meanwhile, this study serves to examine the feasibility of the language resource constructed by the Korean government for deep learning architecture. The research is three-fold. First, we conducted an acceptability judgment testing to verify whether and how the language resource used in this study is indeed trustworthy. The results indicate that the judgments provided in the language resource converge with the judgments of our own experiment well enough. Second, we employed four Korean models such as mBERT, KoBERT, KR-BERT, KorBERT in order to evaluate how the language resource has a potentiality to predict the well-formedness of Korean expressions. The different models yield different results, the reason of which is fully discussed. Third, we made use of an independent test-set for evaluating the deep learning systems. It turns out that the results are still challenging, which implies that the current Korean models may have room for improvement to understand the syntactic phenomena.

키워드

deep learningBERTacceptability judgmentminimal paircorrelation coefficient
제목
최소대립 문장쌍을 활용한 한국어 사전학습모델의 통사 연구 활용 가능성 검증
제목 (타언어)
Verification of Korean Pre-trained Models' Feasibility of Syntactic Research Using Pairwise Sentences
저자
박권식김성태송상헌
발행일
2021
저널명
언어와 정보
25
3
페이지
1 ~ 21