세종 말뭉치에 나타난 한국어 음절의 빈도와 분포

The Distributions and Frequencies of Korean Syllables in Sejong Corpus

초록

The present study aims at building a database of Korean syllable frequencies and distributions as a useful resource that could be consulted by researchers in psycholinguistics and other adjacent disciplines. In doing so, we produced a set of syllable token/type frequency lists by word classes and positions within an eojeol/ headword compiled from the Sejong Corpus containing 15 million eojeols of written texts. The important results include the following: Firstly, the power law was observed, which is characterized by the phenomena that most tokens/types are accounted for by a small number of syllables. Secondly, there was a strong tendency that the token/type frequencies of eojeol/headword syllables decrease as a function of their phonological complexity. Lastly, substantial differences in phonological and morphological aspects were found between the first and second syllables of eojeols/headwords. The database containing 26 different syllable frequency lists can be freely shared via the GitHub repository of one of the authors.

키워드

한국어 음절한국어 음절 빈도한국어 음절 빈도 분포세종 말뭉치Korean syllablesKorean syllable frequenciesdistribution of Korean syllable frequenciesSejong corpus
제목
세종 말뭉치에 나타난 한국어 음절의 빈도와 분포
제목 (타언어)
The Distributions and Frequencies of Korean Syllables in Sejong Corpus
저자
이은하남기춘
발행일
2020
저널명
언어과학연구
92
페이지
79 ~ 130