# Swahili nanoGPT 🌍🤖
**The simplest, fastest way to train Swahili language models**
*Empowering East Africa's digital future with accessible AI technology*
This is a Swahili language model pretrained in a GPT-2 model architecture with 10M parameters, with swahili dataset from OSCAR corpus containing 8.5M characters achieving train loss 1.0362, val loss 1.0904(*These were achieved using a character level encoding, new results after byte pair encoding will be updated after training is finished*).
Tokenization is done using swahilitiktoken a byte pair tokenizer that is trained using pure swahili vocabularies
This is a work in progress so i will keep updating the features in the model
> *Example output:*
> `Sehemu ya ushirika wa tume hiyo, aliyofanya uwezekano wa kulifanikiwa kwa uroho na ya kijaribio vuraa katofauti.
Mkuu wa MES Mgangala akisaidia katika utoaji wa mwito na matoo ya nyoka chini ya upatika Watumishi na Waziri Mkuu akionekana wabunge
wa Shule katika utekelezaji wa barabara ya jima husika kwa mashariki kuwa na kwa mwaka.
Ada hii hoodha hii kutafuta soka kwa wananchi wanafanyiweza kufuata uwezeshaji upinduzi wa kuchukua jana mpambani wa kuwasilia na Rais
wa Jamhuri ya Kimwary.`
## Why Swahili AI Matters
Swahili is spoken by **200+ million people** across East Africa, yet remains underserved in AI development. This repository adapts Karpathy's nanoGPT to:
- 🚀 Democratize Swahili NLP development
- 📈 Build educational/health tools for low-bandwidth regions
- 🔍 Preserve linguistic nuance through lightweight models
- 💡 Enable research on African languages
Contacts
Emily Godfrey Emily | mathematiciangodfrey@outlook.com