Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

amandatheuri/NLP-for-English-Sheng-Swahili

Domaine:

natural language processing

Type de record:

software
Créateur:
ama
Hôte:
This project is a Natural Language Processing (NLP) application that automatically identifies the language of short text inputs. # Language Identification System (NLP Project) ## Overview This project is a Natural Language Processing (NLP) application that automatically identifies the language of short text inputs. It supports multiple languages commonly used in Kenya: * English * Swahili * Sheng * Kikuyu The system was developed as part of an academic project to demonstrate text classification using machine learning techniques. --- ## Objectives * Build a model that can classify short text into different languages * Apply NLP preprocessing and feature extraction techniques * Compare multiple machine learning models * Deploy the model as an interactive application --- ## Technologies Used * Python * Scikit-learn * Pandas / NumPy * Streamlit --- ## Methodology ### 1. Data Collection A balanced dataset was created with text samples from four languages: * English * Swahili * Sheng * Kikuyu Each class contains an equal number of samples. --- ### 2. Data Preprocessing Text data was cleaned using: * Lowercasing * Removal of extra whitespace * Preservation of special characters (important for Kikuyu) --- ### 3. Feature Extraction Text was converted into numerical features using: * TF-IDF Vectorization * Character n-grams (1–5 range) This approach captures language-specific character patterns. --- ### 4. Model Training Two models were trained and compared: * Multinomial Naive Bayes * Logistic Regression --- ### 5. Model Evaluation Performance was evaluated using: * Accuracy * Precision * Recall * F1-score * Confusion Matrix #### Results: * Naive Bayes Accuracy: 100% * Logistic Regression Accuracy: 97.92% Naive Bayes was selected as the final model due to its superior performance. --- ### 6. Deployment The model was deployed using Streamlit, allowing users to: * Enter text * Click a button * Receive an instant language prediction --- ## How to Run the Project ### 1. Clone the repository ```bash git clone github.com …

Visit

github.com

Tasks

language identification

Languages

GikuyuSwahili

Similaires

slilan/English-Swahili-ShengA Sheng Phishing Corpus for Low-Resource Cybersecurity NLPchefjesse/Sheng-Doctor-Kenya-swahili🚧 Free Form Sheng (Swahili-English Code-Switching) Speech in the Mobile Money Domain 🚧Sheng: Rise of a Kenyan Swahili VernacularIMPACT OF SHENG DIALECT ON THE ENGLISH AND SWAHILI TEACHING IN KENYA; A CRITICAL LITERATURE REVIEW

slilan/English-Swahili-Sheng

A Sheng Phishing Corpus for Low-Resource Cybersecurity NLP

 

This dataset, the Sheng-English

chefjesse/Sheng-Doctor-Kenya-swahili

"Sheng Doctor – Kenyan Swahili (Sheng) dataset and tools” # Sheng-Doctor-Kenya-swahili "Sheng Docto

🚧 Free Form Sheng (Swahili-English Code-Switching) Speech in the Mobile Money Domain 🚧

This dataset contains Sheng audio and accurate human-transcript pairs within the mobile money domain

Sheng: Rise of a Kenyan Swahili Vernacular

IMPACT OF SHENG DIALECT ON THE ENGLISH AND SWAHILI TEACHING IN KENYA; A CRITICAL LITERATURE REVIEW

Purpose: Sheng is a blend of two words derived from Kiswahili and English. It is a code created by y