Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Detecting gender bias in Arabic text through word embeddings

Domaine:

natural language processingsocioeconomic

Type de record:

paperdataset
Créateur:
MouAbuElb
Éditeur:
Ame
Éditeur:
CCSDPub
Hôte:avatar
International audience For generations, women have fought to achieve equal rights with those of men. Many historians and social scientists examined this uphill path with a focus on women’s rights and economic status in the West. Other parts of the world, such as the Middle East, remain understudied, with a noticeable shortage in gender-based statistics in the economic arena. According to the sociocognitive theory of critical discourse analysis, social behaviors and norms are reflected by language discourses, which motivates the present study, where we examine gender-based biases in various occupations, as reflected through various textual corpora. Several works in literature have shown that word embedding models can learn biases from the textual data they are trained on, which can propagate societal prejudices that have been implicitly embedded in such text. In our study, we adapt WEAT and Direct Bias quantification tests for Arabic, to examine gender bias with respect to a wide set of occupations as reflected in various Arabic text datasets. These datasets include two Lebanese news archives, Arabic Wikipedia, and electronic newspapers in UAE, Egypt, and Morocco, thus providing different outlooks into female and male engagements in various professions. Our WEAT tests across all datasets indicate that words related to careers, science, and intellectual pursuits are linked to men. In contrast, words related to family and art are associated with women across all datasets. The Direct Bias analysis shows a consistent female gender bias towards professions such as nurse, house cleaner, maid, secretary, and dancer. As the Moroccan News Articles Dataset (MNAD) showed, females were also associated with additional occupations such as researcher, doctor, and professor. Considering that the Arab world remains short on census data exploring gender-based disparities across various professions, our work provides evidence that such stereotypes persist till this day.

Visit

hal.science

Tasks

embeddings

Tags

[INFO.INFO-TT]Computer Science [cs]/Document and Text Processing[INFO.INFO-CL]Computer Science [cs]/Computation and Language [cs.CL][SHS.GENRE]Humanities and Social Sciences/Gender studies

Licenses

https://creativecommons.org/licenses/by/4.0/info:eu-repo/semantics/OpenAccess

Similaires

Detecting Cross-Lingual Plagiarism Using Simulated Word EmbeddingsWhose voice matters? Word embeddings reveal identity bias in news quotesLearning Multilingual Word Embeddings Using Image-Text DataAraWEAT: Multidimensional Analysis of Biases in Arabic Word EmbeddingsDiaLex: A Benchmark for Evaluating Multidialectal Arabic Word EmbeddingsDetecting Emergent Intersectional Biases: Contextualized Word Embeddings Contain a Distribution of Human-like Biases

Detecting Cross-Lingual Plagiarism Using Simulated Word Embeddings

Cross-lingual plagiarism (CLP) occurs when texts written in one language are translated into a diffe

Whose voice matters? Word embeddings reveal identity bias in news quotes

This paper investigates identity bias (gender and race) in the South African news selection and repr

Learning Multilingual Word Embeddings Using Image-Text Data

There has been significant interest recently in learning multilingual word embeddings -- in which se

AraWEAT: Multidimensional Analysis of Biases in Arabic Word Embeddings

Recent work has shown that distributional word vector spaces often encode human biases like sexism o

DiaLex: A Benchmark for Evaluating Multidialectal Arabic Word Embeddings

Word embeddings are a core component of modern natural language processing systems, making the ability to thoroughly evaluate them a vital task. We describe DiaLex, a benchmark for intrinsic evaluation of dialectal Arabic word embedding. DiaLex covers five importan

Detecting Emergent Intersectional Biases: Contextualized Word Embeddings Contain a Distribution of Human-like Biases

With the starting point that implicit human biases are reflected in the statistical regularities of