Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

TunDC: a public benchmark dataset for sentiment analysis and language modeling in the Tunisian dialect

Domain:

natural language processing

Record type:

dataset
Creator:
AhmMou
Publisher:
Institute of Advanced Engineering and Science
Host:
The development of natural language processing (NLP) applications has increasingly focused on dialectal variations of languages. The Tunisian dialect (TD), a widely spoken variant of Arabic, poses unique linguistic challenges due to its lack of standardized writing conventions and influences from multiple languages, including French, Italian, Turkish, and Berber. In this work, we introduce TunDC, a dataset of 20,044 labeled comments designed to advance NLP research on the TD. The dataset covers diverse linguistic forms (Arabic, Latin, and mixed scripts), and each comment was manually annotated for positive or negative sentiment by native speakers, achieving high inter-annotator agreement. To evaluate its effectiveness, we fine-tuned various models on TunDC. The bert-base-arabic-TunDC-mixed model achieved an accuracy of 0.84 and a macro-averaged F1-score of 0.83, demonstrating strong generalization across sentiment categories and writing systems. A stratified data-splitting strategy considering both sentiment and script type further improved accuracy by approximately 8% compared to standard splits. As a publicly available resource, TunDC contributes to the computational linguistics community, fostering advancements in language modeling and applications tailored to the TD.

Visit

doi.org

Tasks

language modelingsentiment analysistext classification

Languages

Arabic, Tunisian SpokenBerber

Licenses

http://creativecommons.org/licenses/by-sa/4.0

Similar

HausaMovieReview: A Benchmark Dataset for Sentiment Analysis in Low-Resource African LanguageCyberDTD: A Multimodal Benchmark Dataset for Cyberbullying Detection in Tunisian DialectSentiment Analysis on Arabic Dialects: A Multi-Dialect BenchmarkTARIC-SLU: A Tunisian Benchmark Dataset for Spoken Language UnderstandingA Sentiment analysis classification for texts in social media:application to the Tunisian dialectazayz/Tunisian-Arabic-Dialect-Sentiment-Analysis

HausaMovieReview: A Benchmark Dataset for Sentiment Analysis in Low-Resource African Language

The development of Natural Language Processing (NLP) tools for low-resource languages is critically

CyberDTD: A Multimodal Benchmark Dataset for Cyberbullying Detection in Tunisian Dialect

International audience Effective detection of cyberbullying requires understanding bo

Sentiment Analysis on Arabic Dialects: A Multi-Dialect Benchmark

TARIC-SLU: A Tunisian Benchmark Dataset for Spoken Language Understanding

A Sentiment analysis classification for texts in social media:application to the Tunisian dialect

Currently, our life is based on information and its analysis. More and more, people communicate, dif

azayz/Tunisian-Arabic-Dialect-Sentiment-Analysis

Sentiment Analysis task on Tunisian and Arabic dialect, data augmentation for NLP and scrapping goog