Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Data Caricatures: On the Representation of African American Language in Pretraining Corpora

Domaine:

natural language processing

Type de record:

paper
Créateur:
DeaVenAnaGri
Hôte:avatar
With a combination of quantitative experiments, human judgments, and qualitative analyses, we evaluate the quantity and quality of African American Language (AAL) representation in 12 predominantly English, open-source pretraining corpora. We specifically focus on the sources, variation, and naturalness of included AAL texts representing the AAL-speaking community. We find that AAL is underrepresented in all evaluated pretraining corpora compared to US demographics, constituting as few as 0.007% and at most 0.18% of documents. We also find that more than 25% of AAL texts in C4 may be perceived as inappropriate for LLMs to generate and to reinforce harmful stereotypes. Finally, we find that most automated filters are more likely to conserve White Mainstream English (WME) texts over AAL in pretraining corpora. ACL 2025

Visit

arxiv.org

Tags

Computation and Language

Similaires

On the Utility of Pretraining Language Models on Synthetic DataEmbracing blackness: The on-screen representation of African American English by white American actorsRevisiting Multilingual Data Mixtures in Language Model PretrainingEffect of African Language Pretraining on XTREME-R Cross-Lingual Transfer PerformanceMultilingual Language Model Pretraining using Machine-translated DataAfrican American students’ representation in S/LI (Robinson & Norton, 2019)

On the Utility of Pretraining Language Models on Synthetic Data

Development of pre-trained language models has predominantly relied on large amounts of datasets. Ho

Embracing blackness: The on-screen representation of African American English by white American actors

Alkalmazott nyelvtudomány, vol. 23. issue 2. ISSN 1587-1061, eISSN 2498-4442

Revisiting Multilingual Data Mixtures in Language Model Pretraining

The impact of different multilingual data mixtures in pretraining large language models (LLMs) has b

Effect of African Language Pretraining on XTREME-R Cross-Lingual Transfer Performance

Recent advances in training multilingual language models on large datasets seem to have shown promis

Multilingual Language Model Pretraining using Machine-translated Data

High-resource languages such as English, enables the pretraining of high-quality large language mode

African American students’ representation in S/LI (Robinson & Norton, 2019)

Purpose: This study aimed to determine if African American students were disproportionat