Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

NabihSamy/Censure_Darija

Domain:

natural language processing

Record type:

softwaremodel
Creator:
Nab
Host:
Censure_Darija is a tool that uses Word2Vec and FastText to detect and censor offensive words in Darija. # Censure Darija An automated offensive content detection and censorship system for Algerian Darija dialect using machine learning and natural language processing techniques. ## 📖 Description This project provides a culturally-adapted content moderation solution for Algerian Darija dialect. It enables automatic detection and censorship of offensive, abusive, or toxic language in Darija text, accounting for the linguistic and cultural specificities of Algeria. ## 🎯 Business Challenges ### Problem Statement - **Expensive Manual Moderation**: Human moderation of Darija content requires significant resources - **Technology Gap**: Existing tools don't understand Algerian dialect nuances - **Digital Content Growth**: Explosion of social platforms and messaging apps in Algeria - **User Protection**: Need to create safe and healthy digital spaces ### Opportunities - **Target Markets**: Social platforms, chat applications, community forums - **Competitive Advantage**: First specialized system for Algerian dialect - **Scalability**: Deployable across various digital platforms ## 🔧 Technical Challenges ### Key Issues - **Under-resourced Language**: Darija has limited available digital resources - **Orthographic Variation**: Writing in both Latin and Arabic scripts with multiple variants - **Code-switching**: Frequent mixing with Standard Arabic, French, and English - **Cultural Expressions**: Specific idioms and insults requiring contextual understanding - **Data Quality**: Collection and annotation of high-quality training data ### Solutions Provided - **Ensemble Models**: Six different ML models (FastText and Word2Vec) for multiple languages - **Advanced Preprocessing**: Normalization and cleaning of multilingual texts - **Linguistic Features**: Extraction of dialect-specific characteristics - **Emoji Detection**: Handling of inappropriate emojis in text ## 🏗️ System Architecture ``` Censure_Darija/ ├── api.py # Main API interface with high-level …

Visit

github.com

Tasks

hate speech detectiontext classification

Languages

Arabic, Algerian Spoken