Logo Lanfrica

MaicyMxtim/habari

Domaine:

natural language processing

Type de record:

projectsoftware
Créateur:
Mai
Hôte:
Swahili news classification: classical baselines vs fine-tuned language models, evaluated with bootstrap confidence intervals # Habari — Swahili news classification with modern and classical NLP Habari (Swahili for "news") is an end-to-end NLP project on low-resource language text classification. It takes the Swahili portion of the MasakhaNEWS dataset and asks a simple question with a rigorous answer: how much does a fine-tuned language model actually buy you over strong classical baselines on a low-resource African language, and how confident can we be in the difference? Most NLP portfolio work reports a single accuracy number on a high-resource English dataset. This project deliberately does neither. Every result below is reported with bootstrap confidence intervals, and every model is compared against a properly tuned classical baseline before any deep learning is used. ## Results | Model | Accuracy | Macro F1 | 95% CI (F1) | |---|---|---|---| | TF-IDF + Logistic Regression | 0.834 | 0.806 | 0.761–0.846 | | Fine-tuned AfriBERTa | 0.874 | 0.857 | 0.818–0.893 | | QLoRA fine-tuned small LLM | _pending_ | _pending_ | _pending_ | Fine-tuning AfriBERTa lifts macro F1 from 0.806 to 0.857. The largest gain lands on the weakest class: entertainment improves from 0.667 to 0.800 F1. The confidence intervals overlap at their edges, which is itself an honest finding on a 476-example test set — the transformer helps, but a single headline accuracy number would overstate how decisively. ## Why Swahili Swahili is spoken by over 100 million people and remains underserved by NLP tooling. Models that look solved on English degrade sharply on low-resource languages, which makes this a more honest test of methods than another English benchmark. The Masakhane community maintains the datasets this project builds on. ## Project structure ``` habari/ src/habari/ data.py # dataset download and preparation evaluate.py # metrics with bootstrap confidence intervals baseline.py # TF-IDF + logistic regression baseline train_transformer.py # AfriBERTa fine-tuning, plain PyTorch loop data/ …