Detailed Information

Cited 0 time in webofscience Cited 0 time in scopus
Metadata Downloads

BERTOEIC: Solving TOEIC Problems Using Simple and Efficient Data Augmentation Techniques with Pretrained Transformer Encoders

Full metadata record
DC Field Value Language
dc.contributor.authorLee, Jeongwoo-
dc.contributor.authorMoon, Hyeonseok-
dc.contributor.authorPark, Chanjun-
dc.contributor.authorSeo, Jaehyung-
dc.contributor.authorEo, Sugyeong-
dc.contributor.authorLim, Heuiseok-
dc.date.accessioned2022-08-12T17:40:59Z-
dc.date.available2022-08-12T17:40:59Z-
dc.date.created2022-08-12-
dc.date.issued2022-07-
dc.identifier.issn2076-3417-
dc.identifier.urihttps://scholar.korea.ac.kr/handle/2021.sw.korea/142933-
dc.description.abstractRecent studies have attempted to understand natural language and infer answers. Machine reading comprehension is one of the representatives, and several related datasets have been opened. However, there are few official open datasets for the Test of English for International Communication (TOEIC), which is widely used for evaluating people's English proficiency, and research for further advancement is not being actively conducted. We consider that the reason why deep learning research for TOEIC is difficult is due to the data scarcity problem, so we therefore propose two data augmentation methods to improve the model in a low resource environment. Considering the attributes of the semantic and grammar problem type in TOEIC, the proposed methods can augment the data similar to the real TOEIC problem by using POS-tagging and Lemmatizing. In addition, we confirmed the importance of understanding semantics and grammar in TOEIC through experiments on each proposed methodology and experiments according to the amount of data. The proposed methods address the data shortage problem of TOEIC and enable an acceptable human-level performance.-
dc.languageEnglish-
dc.language.isoen-
dc.publisherMDPI-
dc.subjectPROFICIENCY-
dc.titleBERTOEIC: Solving TOEIC Problems Using Simple and Efficient Data Augmentation Techniques with Pretrained Transformer Encoders-
dc.typeArticle-
dc.contributor.affiliatedAuthorLim, Heuiseok-
dc.identifier.doi10.3390/app12136686-
dc.identifier.scopusid2-s2.0-85133665261-
dc.identifier.wosid000822108000001-
dc.identifier.bibliographicCitationAPPLIED SCIENCES-BASEL, v.12, no.13-
dc.relation.isPartOfAPPLIED SCIENCES-BASEL-
dc.citation.titleAPPLIED SCIENCES-BASEL-
dc.citation.volume12-
dc.citation.number13-
dc.type.rimsART-
dc.type.docTypeArticle-
dc.description.journalClass1-
dc.description.isOpenAccessY-
dc.description.journalRegisteredClassscie-
dc.description.journalRegisteredClassscopus-
dc.relation.journalResearchAreaChemistry-
dc.relation.journalResearchAreaEngineering-
dc.relation.journalResearchAreaMaterials Science-
dc.relation.journalResearchAreaPhysics-
dc.relation.journalWebOfScienceCategoryChemistry, Multidisciplinary-
dc.relation.journalWebOfScienceCategoryEngineering, Multidisciplinary-
dc.relation.journalWebOfScienceCategoryMaterials Science, Multidisciplinary-
dc.relation.journalWebOfScienceCategoryPhysics, Applied-
dc.subject.keywordPlusPROFICIENCY-
dc.subject.keywordAuthorartificial intelligence-
dc.subject.keywordAuthordeep learning-
dc.subject.keywordAuthornatural language processing-
dc.subject.keywordAuthormachine reading comprehension-
dc.subject.keywordAuthordata augmentation-
Files in This Item
There are no files associated with this item.
Appears in
Collections
Graduate School > Department of Computer Science and Engineering > 1. Journal Articles

qrcode

Items in ScholarWorks are protected by copyright, with all rights reserved, unless otherwise indicated.

Altmetrics

Total Views & Downloads

BROWSE