Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

atlasia/DODa-audio-dataset

Domain:

natural language processing

Record type:

dataset
Creator:
atl
Host:
This dataset consists of 12,743 parallel text and speech samples for Moroccan Darija, including its transcription in both Latin and Arabic scripts and English translations. It was created to support speech recognition, language modeling, and NLP tasks for Moroccan Darija. The dataset was originally sourced from this repository, where it was available as a CSV file containing three columns:

Visit

huggingface.co

Tasks

automatic speech recognitionspeech processing

Languages

Arabic, Algerian SpokenArabic, Moroccan Spoken

Licenses

mit

Similar

atlasia/Moroccan-Darija-Wiki-Audio-Datasetatlasia/scoop-datasetatlasia/Moroccan-Darija-Wiki-Datasetatlasia/Ghibli-style-morocco-datasethaninebou/whisper-mixed-doda-algeriandoda-b/africa-urea-tracker

atlasia/Moroccan-Darija-Wiki-Audio-Dataset

The Moroccan Darija Wiki Audio Dataset consists of 551 parallel text and speech samples of Moroccan

atlasia/scoop-dataset

Language: Arabic (ar) Darija Number of Files: 12 (original sections), Total Number of Rows (Merged):

atlasia/Moroccan-Darija-Wiki-Dataset

The Moroccan Darija Wiki Dataset consists of 10,044 parallel text samples of Moroccan Darija sourced

atlasia/Ghibli-style-morocco-dataset

This dataset contains 10 carefully curated image-text pairs designed to capture Moroccan cultural el

haninebou/whisper-mixed-doda-algerian

doda-b/africa-urea-tracker

# africa-urea-tracker Automated pipeline that fetches urea (HS 310210) import and export data for a