세종 말뭉치에 나타난 한국어 음절의 빈도와 분포

The Distributions and Frequencies of Korean Syllables in Sejong Corpus

초록

The present study aims at building a database of Korean syllable frequencies and distributions as a useful resource that could be consulted by researchers in psycholinguistics and other adjacent disciplines. In doing so, we produced a set of syllable token/type frequency lists by word classes and positions within an eojeol/ headword compiled from the Sejong Corpus containing 15 million eojeols of written texts. The important results include the following: Firstly, the power law was observed, which is characterized by the phenomena that most tokens/types are accounted for by a small number of syllables. Secondly, there was a strong tendency that the token/type frequencies of eojeol/headword syllables decrease as a function of their phonological complexity. Lastly, substantial differences in phonological and morphological aspects were found between the first and second syllables of eojeols/headwords. The database containing 26 different syllable frequency lists can be freely shared via the GitHub repository of one of the authors.

키워드

한국어 음절; 한국어 음절 빈도; 한국어 음절 빈도 분포; 세종 말뭉치; Korean syllables; Korean syllable frequencies; distribution of Korean syllable frequencies; Sejong corpus
제목
세종 말뭉치에 나타난 한국어 음절의 빈도와 분포
제목 (타언어)
The Distributions and Frequencies of Korean Syllables in Sejong Corpus
저자
이은하; 남기춘
발행일
2020
저널명
언어과학연구
호
92
페이지
79 ~ 130