Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

HASSANIYA-DTCD: A new Dataset for Benchmarking Text Classification Tasks on HASSANIYA Dialect

Domain:

natural language processing

Record type:

dataset
Creator:
El
Editor:
Bou
Publisher:
Zenodo
Host:avatar
HASSANIYA-DTCD: A new Dataset for Benchmarking Text Classification Tasks on HASSANIYA dialect is the first Mauritanian dialect dataset called “HASSANIYA” containing 1851 records classified into three categories: positive, negative, and neutral. This dataset was collected using web scraping tools from comments posted on the Facebook platform, and Label Studio was used to annotate each record. For more details, see the README file.

Visit

doi.orgzenodo.org

Tasks

sentiment analysistext classification

Languages

Hassaniyya

Tags

Sentiment AnalysisdialectNatural language processingHASSANIYA

Licenses

Creative Commons Attribution 4.0 Internationalhttps://creativecommons.org/licenses/by/4.0/legalcode

Similar

N-Gram Based HASSANIYA Dialect ClassificationHASSANIYA DatasetDAH (Dataset Hassaniya)Emin009/hassaniya-punctuation-datasetHassaniya Stories OCR DatasetMamadou-Aw/Hassaniya-speech-dataset

N-Gram Based HASSANIYA Dialect Classification

HASSANIYA Dataset

The attached file is the first Mauritanian dialect dataset called “HASSANIYA” containing two thousan

DAH (Dataset Hassaniya)

DAH is a bilingual dataset created to support translation between Hassaniya dialect and English, wit

Emin009/hassaniya-punctuation-dataset

Hassaniya Stories OCR Dataset

Image-to-text OCR pairs extracted from a Hassaniya Arabic stories corpus. Each row couples an image

Mamadou-Aw/Hassaniya-speech-dataset