Data Colonialism and the Suppression of Global South Spanish in LLM Generation: Evidence from GPT-5

초록

This study examines whether implicit interactional cues—specifically, the language of the user's prompt-systematically condition the dialectal variety of Spanish generated by GPT-5, and whether the resulting patterns of activation and suppression reproduce colonial linguistic hierarchies. A corpus of 5,400 outputs across five target varieties (Peninsular, Mexican, Chilean, Argentinian, Peruvian) and six user-language conditions was analyzed through logistic regression, Random Forest predictability analysis, and semantic space modeling. Results reveal a competence-performance gap: GPT-5 generates all varieties under explicit instruction (coefficients up to +1.97) but systematically suppresses Global South varieties under implicit cues, with Argentinian (n=16) and Peruvian (n=21) markers nearly absent. We distinguish a robust competence-performance gap —full generation under explicit instruction, near-absence under implicit cues— from the weaker, partly non-significant gradient of cue sensitivity among implicit conditions; the colonial-default interpretation is grounded primarily in the former. Mexican Spanish achieves the highest lexical predictability (AUC = 0.965) yet no stylistic differentiation in the latent space—tokenistic inclusion without genuine internalization—while Peninsular Spanish shows deeper stylistic coherence despite lower predictability. Interpreted through data colonialism (Couldry & Mejias, 2019), these patterns suggest that RLHF alignment training produces a "colonial default," privileging metropolitan varieties under the guise of communicative safety. The study argues that evaluation must move beyond generative capability to assess willingness to deploy dialectal variation, and that the colonial assumptions embedded in "safety" require critical scrutiny.

키워드

data colonialism/ Large Language Models/ Spanish dialects; 데이터식민주의/거대언어모델/스페인어 방언
제목
Data Colonialism and the Suppression of Global South Spanish in LLM Generation: Evidence from GPT-5
저자
황경진; 이재학
발행일
2026-06
유형
Y
저널명
스페인어문학
호
119
페이지
89 ~ 114