Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

atlasia/Atlaset

Domain:

natural language processing

Record type:

dataset
Creator:
atl
Host:
This dataset is a comprehensive, carefully curated collection of text data specifically for Moroccan Darija, the Arabic dialect spoken in Morocco. It combines various sources to provide a diverse and accurate representation of the language. This dataset was curated by Abdelaziz Bounhar and is particularly suited for tasks such as: Learning word embeddings for Moroccan Darija

Visit

huggingface.co

Tasks

embeddings

Languages

Arabic, Algerian SpokenArabic, Moroccan Spoken

Similar

atlasia/moroccan_darija_domain_classifier_datasetatlasia/darija_bible_alignedatlasia/Darija_LID_Anootation_10katlasia/bible_darija_text_onlyatlasia/moroccan_darija_corpusatlasia/AtlasOCRBench

atlasia/moroccan_darija_domain_classifier_dataset

This dataset is designed for text classification in Moroccan Darija, a dialect spoken in Morocco. It

atlasia/darija_bible_aligned

This dataset contains aligned audio segments from the Moroccan Arabic (Darija) Bible translation, sp

atlasia/Darija_LID_Anootation_10k

This dataset has been created with Argilla. As shown in the sections below, this dataset can be load

atlasia/bible_darija_text_only

atlasia/moroccan_darija_corpus

atlasia/AtlasOCRBench

AtlasOCRBench is a comprehensive evaluation benchmark tailored specifically for Moroccan Darija (Moroccan Arabic dialect) OCR tasks. This dataset was created to measure the real-world performance of OCR models on Darija text, addressing the unique challenges posed