Logo Lanfrica

LuthandoCandlovu/African-Language-NLP-Model-

Domain:

natural language processing

Record type:

modelsoftware
Creator:
Lut
Host:
Sentiment analysis model for isiXhosa and isiZulu using transformers. End-to-end pipeline with preprocessing, training, and Flask API. ``` ╔══════════════════════════════════════════════════════════════════════╗ ║ 🌍 Africa is home to over 2,000 languages. ║ ║ Yet fewer than 1% are well-represented in modern NLP. ║ ║ ║ ║ This project changes that — one sentence at a time. 💚 ║ ╚══════════════════════════════════════════════════════════════════════╝ ``` --- ## 📖 Table of Contents 🗺️ Click to expand navigation | # | Section | Description | |:---:|:---|:---| | 01 | 🌟 Overview | Project introduction & feature matrix | | 02 | ⚡ Quick Start | Get up and running in 5 minutes | | 03 | 🏗 Architecture | System design & data flow diagrams | | 04 | 🔄 Preprocessing Pipeline | Nguni-specific NLP pipeline | | 05 | 🤖 Models | XLM-RoBERTa & AfriBERTa deep dive | | 06 | 🌐 REST API | Endpoints, schemas & examples | | 07 | 📊 Performance | Benchmark results & evaluation metrics | | 08 | 🎥 Live Demo | See it in action | | 09 | 🗂 Project Structure | Full repository layout | | 10 | 💡 Why This Matters | Social & research impact | | 11 | 🤝 Contributing | How to get involved | | 12 | 📜 License | MIT License | --- ## 🌟 Overview This project builds a **state-of-the-art sentiment analysis system** for South African Nguni languages using modern transformer architectures. It is one of the few open-source NLP pipelines specifically designed and optimised for **isiXhosa** and **isiZulu** — spoken by over **12 million people** across Southern Africa. ### ✨ Feature Matrix | 🏷️ Category | ⚙️ Feature | 📌 Details | ✅ Status | |:---:|:---|:---|:---:| | **Languages** | 🗣️ isiXhosa | Full pipeline support | | | | 🗣️ isiZulu | Full pipeline support | | | **Models** | 🧠 XLM-RoBERTa | 125M params · 100 languages | | | | 🧠 AfriBERTa | Purpose-built for Africa | | | **Data Sources** | 🐦 Twitter API | Real-time social data | | | | 📚 NCHLT Corpus | National language corpus | | | **Inference** …