Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

The Arabic Parallel Gender Corpus 2.0: Extensions and Analyses

Domain:

natural language processing

Record type:

paperdataset
Creator:
AlhHabBou
Host:avatar
Gender bias in natural language processing (NLP) applications, particularly machine translation, has been receiving increasing attention. Much of the research on this issue has focused on mitigating gender bias in English NLP models and systems. Addressing the problem in poorly resourced, and/or morphologically rich languages has lagged behind, largely due to the lack of datasets and resources. In this paper, we introduce a new corpus for gender identification and rewriting in contexts involving one or two target users (I and/or You) -- first and second grammatical persons with independent grammatical gender preferences. We focus on Arabic, a gender-marking morphologically rich language. The corpus has multiple parallel components: four combinations of 1st and 2nd person in feminine and masculine grammatical genders, as well as English, and English to Arabic machine translation output. This corpus expands on Habash et al. (2019)'s Arabic Parallel Gender Corpus (APGC v1.0) by adding second person targets as well as increasing the total number of sentences over 6.5 times, reaching over 590K words. Our new dataset will aid the research and development of gender identification, controlled text generation, and post-editing rewrite systems that could be used to personalize NLP applications and provide users with the correct outputs based on their grammatical gender preferences. We make the Arabic Parallel Gender Corpus (APGC v2.0) publicly available.

Visit

arxiv.org

Tags

Computation and Language

Similar

Bamun-French Parallel Corpus 2.0Egyptian Arabic-English Parallel Corpusmbaye930/wolof-arabic-parallel-corpusTunisian Arabic ↔ MSA Parallel CorpusTunisian Arabic → MSA Synthetic Parallel CorpusTunisian Arabic → MSA Synthetic Parallel Corpus

Bamun-French Parallel Corpus 2.0

This dataset is an extended and updated version of the 'Bamun-French Parallel Corpus 1.1' that is pu

Egyptian Arabic-English Parallel Corpus

This dataset is a cleaned and filtered merge of multiple Egyptian Arabic - English parallel corpora,

mbaye930/wolof-arabic-parallel-corpus

Tunisian Arabic ↔ MSA Parallel Corpus

This is an ambitious project to create a high-quality, reproducible parallel corpus for Modern Stand

Tunisian Arabic → MSA Synthetic Parallel Corpus

This dataset is a synthetic parallel corpus of Tunisian Arabic (aeb) and Modern Standard Arabic (arb

Tunisian Arabic → MSA Synthetic Parallel Corpus

This dataset is a synthetic parallel corpus of Tunisian Arabic (aeb) and Modern Standard Arabic (arb