Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Sabri-blm/Building-a-Text-Classification-Model-for-Swahili-News-using-XGBoost

Domain:

natural language processing

Record type:

project
Creator:
Sab
Host:
a repo for Text Classification Model for Swahili News using XGBoost # Swahili News Classification Using XGBoost This repository contains the implementation for building a Swahili news classification model, inspired by the Zindi Swahili News Classification Challenge. The goal is to classify Swahili news articles into predefined categories using machine learning techniques. ## Why Swahili? Swahili is widely spoken in East Africa, with millions of speakers, yet it remains underrepresented in Natural Language Processing (NLP) tools. Developing a Swahili text classification model showcases the richness of the language and contributes to bridging the gap in NLP for underrepresented languages. ## Dataset Overview The dataset consists of Swahili news articles categorized into topics like National, International, Business, Sports, and Entertainment. The distribution of articles is imbalanced, presenting additional challenges during training. - Categories: - `Kitaifa` (National) - `Kimataifa` (International) - `Biashara` (Business) - `Michezo` (Sports) - `Burudani` (Entertainment) - Source: Zindi Swahili News Classification Challenge ## Repository Structure - `code.ipynb`: Main Jupyter Notebook detailing the project workflow, including data preprocessing, exploratory data analysis (EDA), model training, and evaluation. - `train.csv`: File of the training dataset (please download from Zindi, for up to date data). - `stopwords-sw.txt`: Swahili stopwords file used during preprocessing. ## Key Steps ### 1. Data Preprocessing - **Text Normalization:** Lowercasing, removing special characters, numbers, and extra spaces. - **Stopwords Removal:** Utilizing a Swahili stopwords collection from stopwords-iso/stopwords-sw. - **Encoding:** Transforming text data using TF-IDF and categories using label encoding. ### 2. Model Selection and Training We use **XGBoost**, a robust and efficient gradient boosting framework, for multi-class classification. #### Hyperparameter Optimization Bayesian Optimization is employed to fine-tune hyperparameter …

Visit

github.com

Tasks

news classificationtext classificationtopic classification

Languages

Swahili