Already as Native as Native Speakers : Using Fine-Tuned Korean Language Models Based on Korean Language Learner Corpora

초록

This study aimed to evaluate how effectively deep learning-based Korean language models can distinguish between written and spoken sentences produced by Korean language learners and those produced by native Korean speakers. While research on the ability of deep learning models to assess nativelikeness has primarily focused on English, studies on the Korean language remain scarce. To assess the native likeness of Koreans using deep learning language models, this study fine-tuned the KoELECTRA and KLUE-RoBERT a models with corpora representing written and spoken language, as well as other spoken transcription data. The results demonstrated that both language models achieved high accuracy in identifying nativelikeness in both written and spoken language. Specifically, the models achieved approximately 99% and 98% accuracy for written language and 85% and 83% accuracy for spoken language. In addition, we examined how the models evaluated learner sentences containing grammatical and felicity errors. While the overall nativelikeness judgment accuracy of Korean language models is high, there is still a need for improvement in sufficiently refined spoken language data, as well as further investigation through comparative quantitative analysis of native and learner sentences. Nevertheless, this study is significant because it explored nativelikeness detection in both spoken and written Korean, serving as a foundational step in the development of AI-based online tools in foreign language instruction.

키워드

Deep Learning; Nativelikeness Judgment; Grammatical Error; Felicity Error; Korean Learner Corpora
제목
Already as Native as Native Speakers : Using Fine-Tuned Korean Language Models Based on Korean Language Learner Corpora
저자
정윤희; 송상헌
DOI
10.18855/lisoko.2025.50.1.004
발행일
2025-03
저널명
언어
권
50
호
1
페이지
81 ~ 114