Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

lazyAspirations/Sentiment-classification-of-texts-in-Darija

Domaine:

natural language processing
Créateur:
laz
Hôte:
Sentiment Classification of Texts in Darija 🇩🇿🇲🇦 📌 Project Overview This project focuses on the Sentiment Analysis of comments written in Darija (Maghrebi Arabic dialect). The system classifies texts into three categories: Positive, Neutral, and Negative. It leverages Natural Language Processing (NLP) techniques and the K-Nearest Neighbors (KNN) algorithm to handle the unique linguistic challenges of the dialect. Developed by: Aissat Mohamed Moncef 🚀 Workflow & Methodology 1. Data Preprocessing Since Darija is a non-standardized dialect, rigorous cleaning was required: Noise Removal: Stripping punctuation and special characters using re.sub. Normalization: Converting text to lowercase for uniformity. Stop Words Removal: Using a custom list of Maghrebi-specific stop words (e.g., "باش", "واش", "دروك") to filter out non-informative terms. Stemming: Utilizing the ISRIStemmer to reduce Arabic words to their roots. 2. Feature Extraction The text is transformed into numerical data using the TF-IDF (Term Frequency-Inverse Document Frequency) method. Max Features: 5,000. Purpose: This gives higher weight to words that are meaningful for sentiment while downplaying common words. 3. Model Training & Evaluation Algorithm: K-Nearest Neighbors (KNN) with n_neighbors=5 and cosine similarity metric. Data Split: 80% Training / 20% Testing. Metrics: Accuracy, F1-Score, and Confusion Matrix. 📊 Performance Analysis The model evaluates itself using: Confusion Matrix: To visualize where the model confuses "neutral" with "negative" or "positive." Classification Report: Providing Precision, Recall, and F1-score for each class. 🛠️ Tech Stack Language: Python 3 Data Handling: Pandas Machine Learning: Scikit-learn NLP: NLTK (Natural Language Toolkit) Visualization: Matplotlib ⚙️ How to Run Locally 1. Clone the repository: git clone github.com cd Sentiment-classification-of-texts-in-Darija 2. Install …

Visit

github.com

Tasks

sentiment analysistext classification

Languages

Arabic, Algerian SpokenArabic, Libyan SpokenArabic, Moroccan Spoken

Licenses

MIT