Gestion du déséquilibre de classes pour l'analyse de sentiments en dialecte algérien (Darija) — TWIFL corpus — DziriBERT — Master 1 DS & NLP, USDB Blida 1
# Sentiment Analysis on Algerian Dialect (Darija) — Class Imbalance Management
> **Gestion du Desequilibre de Classes pour l'Analyse de Sentiments en Dialecte Algerien**
---
## Project Identity
| Field | Value |
|--------------------|----------------------------------------------------------|
| **University** | USDB Blida 1 — Departement Informatique |
| **Program** | Master 1 Data Science & NLP — Semestre 2 |
| **Module** | Machine Learning |
| **Student** | Abdelaziz Merzoug (solo project) |
| **Supervisor** | Dr. Soraya Cheriguene |
| **Period** | 08 March — 26 April 2026 |
| **Language** | French (comments, analysis) / Python (code) |
---
## Table of Contents
1. What This Project Does
2. Key Concepts — For Beginners
3. Dataset — TWIFL Corpus
4. Model — DziriBERT
5. The Problem: Class Imbalance
6. The Four Strategies Tested
7. Results Summary
8. Project Structure
9. Notebooks — What Each One Does
10. Experimental Protocol
11. Metrics Explained
12. Installation
13. How to Reproduce Results
14. Global Constants — Never Change These
15. Full Results Table
16. Key Findings
17. References
---
## 1. What This Project Does
This project tackles a classic machine learning problem: **what happens when your training data has far more examples of some categories than others?**
We classify Algerian tweets (in "Darija" — a mix of Arabic, French, and Arabizi) into three sentiment categories:
- **Positive** (happy, praise, support)
- **Negative** (criticism, anger, complaint)
- **Neutral** (factual, ambiguous, informational)
The challenge: Positive tweets are almost **4 times more common** than Neutral tweets in our training data. A naive model simply learns to predict "Positive" most o …