A large-scale dataset for korean document-level relation extraction from encyclopedia texts

  • Son, Suhyune
  • Lim, Jungwoo
  • Koo, Seonmin
  • Kim, Jinsung
  • Kim, Younghoon
  • ... Lim, Heuiseok
  • 외 2명
Citations

WEB OF SCIENCE

0
Citations

SCOPUS

1

초록

Document-level relation extraction (RE) aims to predict the relational facts between two given entities from a document. Unlike widespread research on document-level RE in English, Korean document-level RE research is still at the very beginning due to the absence of a dataset. To accelerate the studies, we present TREK (Toward Document-Level Relation Extraction in Korean) dataset constructed from Korean encyclopedia documents written by the domain experts. We provide detailed statistical analyses for our large-scale dataset and human evaluation results suggest the assured quality of TREK . Also, we introduce the document-level RE model that considers the named entity-type while considering the Korean language's properties. In the experiments, we demonstrate that our proposed model outperforms the baselines and conduct qualitative analysis.

키워드

Natural Language ProcessingInformation ExtractionDocument-level Relation ExtractionKorean Relation ExtractionENTITY
제목
A large-scale dataset for korean document-level relation extraction from encyclopedia texts
저자
Son, SuhyuneLim, JungwooKoo, SeonminKim, JinsungKim, YounghoonLim, YoungsikHyun, DongseokLim, Heuiseok
DOI
10.1007/s10489-024-05605-9
발행일
2024-07-02
유형
Article
저널명
Applied Intelligence
54
17-18
페이지
8681 ~ 8701