The Moroccan Darija Wiki Audio Dataset consists of 551 parallel text and speech samples of Moroccan Darija sourced from Wikipedia Darija . This dataset is designed to support speech recognition, language modeling, and various NLP tasks for Moroccan Darija.
The data was scraped from Wikipedia (ary) using the WikiScraper tool.
Data Preprocessing