This dataset is a cleaned corpus of Tunisian Arabic dialect text, aggregated from multiple public sources on Hugging Face. It is designed for Continual Pretraining (CPT) and general NLP research.
A dedicated preprocessing pipeline was applied to:
Normalize text
Remove noise and artifacts
Filter non-Arabic content
Ensure higher overall data quality