Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

TunDC: a public benchmark dataset for sentiment analysis and language modeling in the Tunisian dialect

Domaine:

natural language processing

Type de record:

dataset
Créateur:
AhmMou
Éditeur:
Institute of Advanced Engineering and Science
Hôte:
The development of natural language processing (NLP) applications has increasingly focused on dialectal variations of languages. The Tunisian dialect (TD), a widely spoken variant of Arabic, poses unique linguistic challenges due to its lack of standardized writing conventions and influences from multiple languages, including French, Italian, Turkish, and Berber. In this work, we introduce TunDC, a dataset of 20,044 labeled comments designed to advance NLP research on the TD. The dataset covers diverse linguistic forms (Arabic, Latin, and mixed scripts), and each comment was manually annotated for positive or negative sentiment by native speakers, achieving high inter-annotator agreement. To evaluate its effectiveness, we fine-tuned various models on TunDC. The bert-base-arabic-TunDC-mixed model achieved an accuracy of 0.84 and a macro-averaged F1-score of 0.83, demonstrating strong generalization across sentiment categories and writing systems. A stratified data-splitting strategy considering both sentiment and script type further improved accuracy by approximately 8% compared to standard splits. As a publicly available resource, TunDC contributes to the computational linguistics community, fostering advancements in language modeling and applications tailored to the TD.

Visit

doi.org

Tasks

language modelingsentiment analysistext classification

Languages

Arabic, Tunisian SpokenBerber

Licenses

http://creativecommons.org/licenses/by-sa/4.0

Similaires

HausaMovieReview: A Benchmark Dataset for Sentiment Analysis in Low-Resource African LanguageCyberDTD: A Multimodal Benchmark Dataset for Cyberbullying Detection in Tunisian DialectSentiment Analysis on Arabic Dialects: A Multi-Dialect BenchmarkTARIC-SLU: A Tunisian Benchmark Dataset for Spoken Language UnderstandingA Sentiment analysis classification for texts in social media:application to the Tunisian dialectazayz/Tunisian-Arabic-Dialect-Sentiment-Analysis

HausaMovieReview: A Benchmark Dataset for Sentiment Analysis in Low-Resource African Language

The development of Natural Language Processing (NLP) tools for low-resource languages is critically

CyberDTD: A Multimodal Benchmark Dataset for Cyberbullying Detection in Tunisian Dialect

International audience Effective detection of cyberbullying requires understanding bo

Sentiment Analysis on Arabic Dialects: A Multi-Dialect Benchmark

TARIC-SLU: A Tunisian Benchmark Dataset for Spoken Language Understanding

A Sentiment analysis classification for texts in social media:application to the Tunisian dialect

Currently, our life is based on information and its analysis. More and more, people communicate, dif

azayz/Tunisian-Arabic-Dialect-Sentiment-Analysis

Sentiment Analysis task on Tunisian and Arabic dialect, data augmentation for NLP and scrapping goog