Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

drelhaj/Arabic-Dialects

Domaine:

natural language processing

Type de record:

dataset
Créateur:
dre
Hôte:
The Arabic Dialects Dataset is a specialised corpus designed for automatic dialect identification, with a focus on the linguistic phenomena of bivalency and written code-switching between major Arabic dialects and Modern Standard Arabic (MSA).It covers five varieties: EGY – Egyptian Arabic GLF – Gulf Arabic LAV – Levantine Arabic NOR – North African / Tunisian Arabic MSA – Modern Standard Arabic

Visit

huggingface.co

Tasks

code switchinglanguage identification

Languages

Arabic, Tunisian Spoken

Tags

arabicdialect-identificationmultilingualsociolinguisticscode-switchingbivalency

Licenses

cc-by-4.0

Similaires

The Maghrebi dialects of Arabicmalek-hedhli/Arabic-Dialects-DatasetContact-induced grammaticalization between Arabic dialectsArt. XIV.—Dialects of Colloquial ArabicDiacritization of Maghrebi Arabic Sub-DialectsQCRI Arabic Dialects Identification (QADI) Corpus

The Maghrebi dialects of Arabic

This chapter analyses synchronically and diachronically the Maghrebi Arabic dialects spoken in North

malek-hedhli/Arabic-Dialects-Dataset

Arabic tweets, posts and comments in 5 Arabic dialects represent as follows 4834 Tunisian 4834 MGH 4

Contact-induced grammaticalization between Arabic dialects

This chapter describes the phenomenon of contact-induced grammaticalization\linebreak between Arabic dialects and its significance in accounting for the development of future tense markers across modern Arabic varieties. After an introduction to theoretical aspects

Art. XIV.—Dialects of Colloquial Arabic

The Arabic language is commonly spoken throughout a very large area of the old hemisphere. It is the

Diacritization of Maghrebi Arabic Sub-Dialects

Diacritization process attempt to restore the short vowels in Arabic written text; which typically a

QCRI Arabic Dialects Identification (QADI) Corpus

QCRI Arabic Dialects Identification (QADI) is a Country-level Arabic dialects identification (DI) dataset. It provides a collection for benchmarking DI task.The dataset contains 540,590 tweets from 18 Arab countries.