Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Arabic Dialect Corpus

Domain:

natural language processing

Record type:

dataset
Creator:
dat
Host:
A comprehensive collection of Arabic dialectal text, standardized for Natural Language Processing (NLP) model training, evaluation, and linguistic analysis. This corpus has been meticulously processed to ensure high-quality tokenization and consistent metadata. Metric Value Total Records 127,180 Total Tokens 5,802,324 Average Tokens per Record 45.62 Dialect Categories 5

Visit

huggingface.co

Tasks

language modeling

Languages

Arabic, Egyptian SpokenArabic, Moroccan Spoken

Tags

arabicdialectsnlpspeech-to-texttranscriptiontext-classificationlinguisticscorpusegyptiangulf+4

Licenses

mit

Similar

Arabic Dialect Corpus (Egyptian & Saudi)Arabic Dialect Corpus (Egyptian & Saudi)Machine Translation Experiments on PADIC: A Parallel Arabic DIalect CorpusTarab: A Multi-Dialect Corpus of Arabic Lyrics and PoetryARCADE: A City-Scale Corpus for Fine-Grained Arabic Dialect TaggingArabic dialect identification: An in-depth error analysis on the MADAR parallel corpus

Arabic Dialect Corpus (Egyptian & Saudi)

This dataset contains 150K+ natural, informal Arabic text samples scraped from high-engagement YouTu

Arabic Dialect Corpus (Egyptian & Saudi)

This dataset contains 150K+ natural, informal Arabic text samples scraped from high-engagement YouTu

Machine Translation Experiments on PADIC: A Parallel Arabic DIalect Corpus

We present in this paper PADIC, a Parallel Arabic DIalect Corpus we built from scratch, then we conducted experiments on cross-dialect Arabic machine translation. PADIC is composed of dialects from both the Maghreb and the Middle-East. Each dialect has been aligned

Tarab: A Multi-Dialect Corpus of Arabic Lyrics and Poetry

We introduce the Tarab Corpus, a large-scale cultural and linguistic resource that brings together A

ARCADE: A City-Scale Corpus for Fine-Grained Arabic Dialect Tagging

The Arabic language is characterized by a rich tapestry of regional dialects that differ substantial

Arabic dialect identification: An in-depth error analysis on the MADAR parallel corpus