π²π¦ Darija Sentiment Analysis
ΨͺΨΩΩΩ Ψ§ΩΩ
Ψ΄Ψ§ΨΉΨ± Ψ¨Ψ§ΩΨ―Ψ§Ψ±Ψ¬Ψ© Ψ§ΩΩ
ΨΊΨ±Ψ¨ΩΨ©
First open-source, end-to-end NLP pipeline for Moroccan Arabic (Darija)
From raw web scraping β preprocessing β model training β live inference
> *"40 million people speak Darija. Until now, AI couldn't understand a single word of their sentiment."*
---
## π Table of Contents
- The Problem
- What We Built
- Live Demo
- Dataset
- Data Analysis
- System Architecture
- Preprocessing Pipeline
- Models & Results
- Project Structure
- Quick Start
- Research Findings
- Roadmap
- Academic Context
---
## π¨ The Problem No One Solved
**Moroccan Arabic (Darija)** is spoken by over **40 million people** daily. It is a living, evolving language β a unique fusion of:
- π€ **Classical Arabic** β grammatical backbone
- π«π· **French** β embedded in everyday speech
- ποΈ **Tamazight (Berber)** β indigenous vocabulary
- πͺπΈ **Spanish** β in northern regions
Despite this scale, **Darija is one of the most NLP-neglected dialects in the world.**
| Challenge | Impact |
|---|---|
| No standard spelling | Same word = 5+ forms |
| Mixed scripts | Arabic + Latin + digits |
| No labeled datasets | Training impossible |
| No preprocessing tools | Zero libraries exist |
| Dialectal variation | Region-to-region differences |
**The consequence:** Moroccan companies, researchers, and public institutions cannot perform automated sentiment analysis on the content their users generate every day β social media, reviews, news comments, customer feedback.
**This project is the first complete answer to that problem.**
---
## π― What We Built
An **end-to-end NLP pipeline** covering every stage from raw data to live inference:
```
βββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
β FULL PROJECT PIPELINE β
β β
β π Web Scraping β π§Ή Preprocessing β π·οΈ Labelin β¦