상세 보기
대규모 신문 기사의 자동 키워드 추출과 분석 -t-점수를 이용하여-
- 김일환;
- 이도길
초록
Kim, Ilhwan & Lee, Do-Gil. 2011. 11. Automatic Keyword Extraction and Analysis from the Large Scale Newspaper Corpus Based on t-score. Korean Linguistics 53,145-194. As the type and size of documents radically increased in recent years, how to automatically extract proper keywords from those documents has also been important. This paper aims to propose an automatic method to extract keywords and to analyze their characteristics. The keywords are extracted from Trends 21 corpus, a collection of four major Korean daily newspapers issued from the year 2000 to 2009. We introduce t-score to measure the keywordness. The keywords were extracted from two aspects i.e. year and topic. We present the top 100 keywords for 6 topics and 10years. Also, to verify whether these keywords can be representatives of the texts, we compared them with the headline news of 2009. The two main contributions of this work are as follows: 1) this study can present keywords which are automatically extracted from large scaled corpora without any human intervention by the verifiable and objective method and 2) this study analyzed the characteristics of the keywords by topic and year.