Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

M20Jay/kiswahili-nlp

Domaine:

natural language processing

Type de record:

model
Créateur:
M20
Hôte:
Kiswahili environmental text classifier using mBERT — classifying East African language text by UNEP Strategic Objective. Built by Martin James Ng'ang'a | github.com # Kiswahili NLP — Environmental Text Classifier --- ## The Problem Over 200 million East Africans speak Kiswahili as a first or second language. Yet most AI systems are built primarily in English — leaving indigenous communities unable to contribute environmental knowledge in their own language. Community observations about deforestation, pollution, and climate change remain invisible to global monitoring systems because the AI cannot process them. This project addresses that gap. --- ## What This Does A Kiswahili environmental text classifier that: - Classifies Kiswahili text by UNEP Strategic Objective - Uses mBERT — multilingual BERT trained on 104 languages including Kiswahili - Tracks every experiment with MLflow - Connects African language knowledge to global environmental monitoring ## Where This Fits A Kiswahili classification model, even fully built, doesn't create value sitting in a repository — it needs to reach the systems where environmental data actually gets aggregated: - **UNEP monitoring integration** — classified Kiswahili observations, mapped to Strategic Objectives, are exactly the kind of structured input global environmental monitoring dashboards currently can't ingest from non-English sources. - **Citizen science platforms** — community reporting apps (deforestation alerts, pollution reports) currently require English input to be processed by most downstream systems; this closes the gap at the language layer, not the reporting layer. - **Cross-language research aggregation** — researchers studying environmental trends across East Africa currently can't systematically include Kiswahili-language community knowledge in quantitative analysis; structured classification is what makes that possible. This currently runs on zero-shot classification (no fine-tuning yet) — genuinely functional, but the accuracy ceiling of an off-the-shelf model. The next step is fine-tuning mBERT on a real Kiswahili environmental corpus, which is what would m …

Visit

github.com

Languages

SwahiliSwahili, CoastalSwahili, Congo

Similaires

TimothMutua/NLP-PROJECT-KISWAHILI-SENTIMENT-ANALYZERKinyarwanda-NLP/NLPkiswahili-rahisi/kiswahili-rahisimsgyi/NLPMoroccan NLPtwamaa/NLP

TimothMutua/NLP-PROJECT-KISWAHILI-SENTIMENT-ANALYZER

NLP PROJECT KISWAHILI SENTIMENT ANALYZER

Kinyarwanda-NLP/NLP

NLP in Kinyarwanda # NLP NLP in Kinyarwanda

kiswahili-rahisi/kiswahili-rahisi

Audio Pronunciation Guide for Kiswahili Rahisi Book

msgyi/NLP

Natural Languge Processing of Asante Twi Language Ghanaian Asante Twi is the most widely spoken ind

Moroccan NLP

Development and evaluation of NLP models for Moroccan Darija and Arabic, covering code-switching det

twamaa/NLP

LSTM, CNN and Transformer Deep Learning Models for Text Classification Tasks in Swahili # Swahili N