This dataset contains text in Algerian Darija, collected from a variety of sources including existing datasets on Hugging Face, web scraping, and YouTube transcript APIs.
The train split consists more then 2k rows of uncleaned text data.
The v1 split consists more than 170k rows of split and partially cleaned text.
The text data was gathered from:
Hugging Face Datasets: Pre-existing datasets relevant to Algerian Darija.