Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Paul-mwaura/Zindi-Sentiment-Analysis_Tunisian-Arabizi

Domain:

natural language processing

Record type:

dataset
Creator:
Pau
Host:
Can you classify sentiment in the Tunisian Arabizi dialect? # Zindi: Sentiment-Analysis_Tunisian-Arabizi Can you classify sentiment in the Tunisian Arabizi dialect? ## Problem Statement On social media, Arabic speakers tend to express themselves in their own local dialect. To do so, Tunisians use ‘Tunisian Arabizi’, where the Latin alphabet is supplemented with numbers. However, annotated datasets for Arabizi are limited; in fact, this challenge uses the only known Tunisian Arabizi dataset in existence. Sentiment analysis relies on multiple word senses and cultural knowledge, and can be influenced by age, gender and socio-economic status.For this task, we have collected and annotated sentences from different social media platforms. The objective of this challenge is to, given a sentence, classify whether the sentence is of positive, negative, or neutral sentiment. For messages conveying both a positive and negative sentiment, whichever is the stronger sentiment should be chosen. Predict if the text would be considered positive, negative, or neutral (for an average user). This is a binary task. Such solutions could be used by banking, insurance companies, or social media influencers to better understand and interpret a product’s audience and their reactions. ## Data Understanding TUNIZI is the first 100% Tunisian Arabizi sentiment analysis dataset, developed as part of AI4D’s ongoing NLP project for African languages. Tunisian Arabizi is the representation of the Tunisian dialect written in Latin characters and numbers rather than Arabic letters. iCompass gathered comments from social media platforms that express sentiment about popular topics. For this purpose, we extracted 100k comments using public streaming APIs. Tunizi was preprocessed by removing links, emoji symbols, and punctuations. The collected comments were manually annotated using an overall polarity: positive (1), negative (-1) and neutral (0). The annotators were diverse in gender, age and social background. ### Variable definition: text_id: Unique ide …

Visit

github.com

Tasks

sentiment analysistext classification

Languages

Arabic, Tunisian Spoken

Licenses

MIT

Similar

Paul-mwaura/Tanzania-Tourism-Hackathon-Zindijkrajanowski/darija-arabizi-sentimentNazarioR9/arabizi-sentiment-analysisnesrinewagaa/Tunisian-Arabizi-Sentiment-Analysis

Paul-mwaura/Tanzania-Tourism-Hackathon-Zindi

# Tanzania-Tourism-Hackathon-Zindi ## Problem Statement The Tanzanian tourism sector plays a signif

jkrajanowski/darija-arabizi-sentiment

# Sentiment Analysis of Moroccan Darija Arabizi Core code and re-annotated dataset for a three-clas

NazarioR9/arabizi-sentiment-analysis

Tunisian Arabizi dialect sentiment analysis # AI4D iCompass Social Media Sentiment Analysis for Tun

nesrinewagaa/Tunisian-Arabizi-Sentiment-Analysis

# Tunisian Arabizi Dialect Data - Sentiment Analysis ## Project Overview This project focuses on S