Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Syrinesmati/tunisian-dialect-corpus

Domain:

natural language processing

Record type:

dataset
Creator:
Syr
Host:
This dataset is a cleaned corpus of Tunisian Arabic dialect text, aggregated from multiple public sources on Hugging Face. It is designed for Continual Pretraining (CPT) and general NLP research. A dedicated preprocessing pipeline was applied to: Normalize text Remove noise and artifacts Filter non-Arabic content Ensure higher overall data quality

Visit

huggingface.co

Tasks

language modeling

Languages

Arabic, Tunisian Spoken

Licenses

apache-2.0

Similar

Syrinesmati/tunisian-question-response-datasetNaim Mhedhbi Tunisian Dialect Corpus v1Naim Mhedhbi Tunisian Dialect Corpus v0Tunisian Dialect Speech Corpus: Construction and Emotion AnnotationArabic Dialect CorpusHabibBelguith44/Llama3-Tunisian-Dialect

Syrinesmati/tunisian-question-response-dataset

A supervised fine-tuning (SFT) dataset of ~31,000 question-answer pairs in Tunisian Arabic dialect (

Naim Mhedhbi Tunisian Dialect Corpus v1

Tunisian dialect corpus for text classification

Naim Mhedhbi Tunisian Dialect Corpus v0

Tunsian dialect corpus for sentiment analysis

Tunisian Dialect Speech Corpus: Construction and Emotion Annotation

Arabic Dialect Corpus

A comprehensive collection of Arabic dialectal text, standardized for Natural Language Processing (N

HabibBelguith44/Llama3-Tunisian-Dialect