상세 보기
초록
This paper aims to demonstrate the applicability of topic modeling, which can organize and summarize large archives of texts, from a corpus-linguistic perspective. To do this, we investigate thematic structures in the Brown Corpus uncovered by an R package which implements topic modeling based on LDA (latent Dirichlet allocation), and use statistical techniques such as comparison cloud, principal component analysis and phylogenetic tree to analyze and visualize the results effectively. This paper shows (i) that the Brown Corpus has a core thematic structure which is divided into texts representing the tendency of past tense and spoken language and texts representing the tendency of present tense and written language, (ii) that the former texts are mainly about women, home, and battle, and the latter texts are primarily related to humanities, society and the economy, and (iii) that the linguistic texts reveal the interdisciplinary nature related to mathematics and engineering, as well as humanities and social sciences.
키워드
- 제목
- 토픽모델링을 이용한 코퍼스의 주제구조 탐색
- 제목 (타언어)
- Exploring the Thematic Structure in Corpora with Topic Modeling
- 저자
- 홍정하; 최재웅
- 발행일
- 2017
- 저널명
- 언어와 정보 사회
- 권
- 30
- 페이지
- 239 ~ 276