Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Egyptian Arabic-English Parallel Corpus

Domain:

natural language processing

Record type:

dataset
Creator:
Ibr
Host:
This dataset is a cleaned and filtered merge of multiple Egyptian Arabic - English parallel corpora, containing ~27,000 aligned sentence pairs. It’s designed for researchers and developers working on machine translation, speech translation, and other NLP tasks involving Egyptian Arabic and English. Sources 📚

Visit

huggingface.co

Tasks

machine translation

Languages

Arabic, Egyptian Spoken

Licenses

mit

Similar

Akan–English Maternal Health Parallel Corpus Akan–English Maternal Health Parallel CorpusSomali-English Parallel CorpusNupe-English parallel corpusAmharic-English Parallel Corpusluganda-english-parallel-corpusEnglish Hausa Parallel Corpus

Akan–English Maternal Health Parallel Corpus Akan–English Maternal Health Parallel Corpus

This dataset contains a curated bilingual parallel corpus developed to support domain-specific neura

Somali-English Parallel Corpus

This dataset contains high-quality parallel sentence pairs, multi-sentence alignments, and paragraph

Nupe-English parallel corpus

This is the first ever Nupe - English Parallel Corpus and Nupe Monolingual Corpora curated from diverse sources including poems,idioms, proverbs, religpoius text etc. The aim of this data collection is to make available a cultural-aware Nupe-english corpus for NLP

Amharic-English Parallel Corpus

This corpus consists of 145,820 Amharic-English parallel sentences (segments) from various sources. This corpus is larger in size than previously compiled corpora. It is released for research purposes and can be used to train or support Amharic-English machine tran

luganda-english-parallel-corpus

English Hausa Parallel Corpus

This English–Hausa Parallel Corpus is a curated bilingual dataset of 5,000 aligned sentence pairs, t