Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Probing Gender Bias in Masked Language Models for Low-Web Data Languages

Domain:

natural language processing

Record type:

paper
Creator:
AssBal
Publisher:
Und
Host:avatar
Low-resourced languages are increasingly included in large multilingual models. While including more languages in pretrained models is a sign of progress, large models still underperform on low-resourced languages. In prioritizing scale over effective processing, we risk 1) deploying language technologies that misrepresent these languages and 2) amplifying gender biases embedded in training corpora. In this paper, we investigate how masked language models encode gender for three low-web-data languages, Afan Oromo, Amharic, and Tigrinya, and how these representations shift after continued pretraining on NLLB data. Using a controlled cloze-style probing setup, we examine prediction patterns. Our findings show consistent gender asymmetries and predictions aligned with stereotypical adjectives and occupations. After continued pretraining, we find that male-gendered predictions reach up to 68% in Amharic, while neutral predictions exceed 60% in Afan Oromo. Our work shows that expanding training data does not guarantee balanced gender representations without careful consideration in data curation.

Visit

doi.orgunderline.io

Languages

AmharicOromoOromo, Borana-Arsi-GujiOromo, EasternOromo, West CentralTigrigna

Tags

Computational LinguisticsNatural Language ProcessingArtificial Intelligence