BERT Learns More than Word Frequency Information: A Case Study of Do-Be Constructions

BERT Learns More than Word Frequency Information: A Case Study of Do-Be Constructions

초록

This study aims to understand BERT’s linguistic ability using naturally occurring data. In particular, the study collected marginal language data, such as what we do is create Frankenstein, which is referred to as a Do-Be construction (DBC) (Flickinger & Wasow, 2013). Using web corpora, the study first collected 17,737 instances of the DBC across text genres and English dialects. The corpus analysis supports the idea that DBC is a computationally challenging phenomenon for data-driven language systems due to its statistical sparsity and linguistic complexity. With manual annotations of DBCs, the study designed two computational prediction tasks: subject―verb agreement and synonym substitution tasks, based on the introspective judgment of linguists. The study found that BERT is hugely sensitive to linguistic acceptability of grammatical forms and felicitous words in the prediction tasks, even though the target phenomenon is rarely observed in corpus data. These results show that the neural language model, BERT, can learn abstract linguistic properties beyond surface frequency information.

키워드

Do-Be constructionneural language modelweb corporaagreement attractionsynonym substitution
제목
BERT Learns More than Word Frequency Information: A Case Study of Do-Be Constructions
제목 (타언어)
BERT Learns More than Word Frequency Information: A Case Study of Do-Be Constructions
저자
신운섭송상헌
DOI
10.18855/lisoko.2022.47.3.004
발행일
2022
저널명
언어
47
3
페이지
467 ~ 489