The Korean Coronavirus Corpus: A Large-Scale Analysis Using Computational Skills

이규민; 송상헌

doi:10.14353/sjk.2022.30.3.08

Detailed Information

Cited 0 time in webofscience

Cited 0 time in scopus

Metadata Downloads

The Korean Coronavirus Corpus: A Large-Scale Analysis Using Computational Skills

Full metadata record

DC Field	Value	Language
dc.contributor.author	이규민	-
dc.contributor.author	송상헌	-
dc.date.accessioned	2022-10-06T15:41:51Z	-
dc.date.available	2022-10-06T15:41:51Z	-
dc.date.created	2022-10-06	-
dc.date.issued	2022	-
dc.identifier.issn	1226-4822	-
dc.identifier.uri	https://scholar.korea.ac.kr/handle/2021.sw.korea/144133	-
dc.description.abstract	Despite the massive impact of COVID-19 on society, beyond the numbers of confirmed cases and deaths, there remains a lack of large-scale data depicting the effects of the virus on the society of the Republic of Korea. To fill this gap, we collected 1.822 million news articles with more than 1 billion morphemes from January 2020 to June 2022, creating a Korean version of the Coronavirus Corpus. This corpus is introduced in the current study. In addition, to demonstrate how such massive corpus can be utilized, we conducted information theoretical analyses to see how the stance of the press media on topics such as vaccines and social distancing affected the COVID-19 situation in the Republic of Korea. Specifically, we utilized several computational linguistic skills including concordance building and sentiment analysis through both traditional and machine learning techniques and measured the transfer entropy to estimate the impact with information theory. The results suggest that the overall impact of the press media on the society was minimal to non-existent.	-
dc.language	English	-
dc.language.iso	en	-
dc.publisher	한국사회언어학회	-
dc.title	The Korean Coronavirus Corpus: A Large-Scale Analysis Using Computational Skills	-
dc.title.alternative	The Korean Coronavirus Corpus: A Large-Scale Analysis Using Computational Skills	-
dc.type	Article	-
dc.contributor.affiliatedAuthor	송상헌	-
dc.identifier.doi	10.14353/sjk.2022.30.3.08	-
dc.identifier.bibliographicCitation	사회언어학, v.30, no.3, pp.213 - 243	-
dc.relation.isPartOf	사회언어학	-
dc.citation.title	사회언어학	-
dc.citation.volume	30	-
dc.citation.number	3	-
dc.citation.startPage	213	-
dc.citation.endPage	243	-
dc.type.rims	ART	-
dc.identifier.kciid	ART002881708	-
dc.description.journalClass	2	-
dc.description.journalRegisteredClass	kci	-
dc.subject.keywordAuthor	COVID-19	-
dc.subject.keywordAuthor	media	-
dc.subject.keywordAuthor	Republic of Korea	-
dc.subject.keywordAuthor	corpus	-
dc.subject.keywordAuthor	computational linguistics	-
dc.subject.keywordAuthor	sentiment analysis	-
dc.subject.keywordAuthor	diachronic analysis	-

Files in This Item: There are no files associated with this item.

Appears in Collections: College of Liberal Arts > Department of Linguistics > 1. Journal Articles

Show simple item record

qrcode

Altmetrics

Total Views & Downloads

STATISTICS: Total View :9,533,268; Today View :26,212

RSS_1.0 RSS_2.0 ATOM_1.0

(02841) 서울특별시 성북구 안암로 14502-3290-1114

Certain data included herein are derived from the © Web of Science of Clarivate Analytics. All rights reserved.
You may not copy or re-distribute this material in whole or in part without the prior written consent of Clarivate Analytics.

Detailed Information

Altmetrics

Total Views & Downloads

BROWSE