Profane or Not: Improving Korean Profane Detection using Deep Learning

Citations

WEB OF SCIENCE

6
Citations

SCOPUS

7

초록

Abusive behaviors have become a common issue in many online social media platforms. Profanity is common form of abusive behavior in online. Social media platforms operate the filtering system using popular profanity words lists, but this method has drawbacks that it can be bypassed using an altered form and it can detect normal sentences as profanity. Especially in Korean language, the syllable is composed of graphemes and words are composed of multiple syllables, it can be decomposed into graphemes without impairing the transmission of meaning, and the form of a profane word can be seen as a different meaning in a sentence. This work focuses on the problem of filtering system mis-detecting normal phrases with profane phrases. For that, we proposed the deep learning-based framework including grapheme and syllable separation-based word embedding and appropriate CNN structure. The proposed model was evaluated on the chatting contents from the one of the famous online games in South Korea and generated 90.4% accuracy.

키워드

Profanitydeep learningconvolutional neural networktext miningnatural language processing
제목
Profane or Not: Improving Korean Profane Detection using Deep Learning
저자
Woo, JiyoungPark, Sung HeeKim, Huy Kang
DOI
10.3837/tiis.2022.01.017
발행일
2022-01-31
유형
Article
저널명
KSII Transactions on Internet and Information Systems
16
1
페이지
305 ~ 318