언어 자료를 활용한 한국어 복합명사 구조 분석

  • 김동성

초록

This paper introduces an analysis on the structure of Korean complex nominals. The analysis attracts both theoretical linguistics and language processing related areas, such as information retrieval and speech synthesis. Our approach has three stages. First, we identify endocentric data and exocentric data, using human intuition. Since exocentric data does not have the internal structure, we do not consider exocentric data. Second, we do the bracketing experiment on the endocentric data for representing a hierarchical structure of constituent parts, using statistical collocation measurements based on the 10 million Sejong corpus. The last stage is composed of several processes to figure out head-modifier or predicate-argument relations, using argument structure and selection restriction specified in the Sejong electronic dictionary. Our method is based on not only the corpus-based materials but also linguistic knowledge (with intuition-based judgement). The importance of our approach is to show how to use language resources to utilize linguistic knowledges in analyzing linguistic data.

키워드

MorphologyCollocationComplex NominalsComputational LinguisticsCorpus StatisticsComputational MorphologyEndocentricity/ExocentricityBracketingMorphologyCollocationComplex NominalsComputational LinguisticsCorpus StatisticsComputational MorphologyEndocentricity/ExocentricityBracketing
제목
언어 자료를 활용한 한국어 복합명사 구조 분석
저자
김동성
DOI
10.24303/lakdoi.2011.19.3.129
발행일
2011
저널명
언어학
19
3
페이지
129 ~ 150