An open corpus for the endangered and underrepresented languages of East Africa
# loba.
**Languages found, never lost.**
An open corpus for the endangered and underrepresented languages of East Africa — built by the speakers, for the speakers and the tools that will serve them.
## What is Loba?
Loba is a community-built, openly licensed dataset of words, phrases, proverbs, and sentences in the languages of East Africa. It starts with **Dholuo** and is designed to grow.
The corpus exists so that developers, researchers, and educators can build spell checkers, translation tools, dictionaries, and AI assistants that actually work for Luo speakers — and eventually speakers of Kikuyu, Kamba, Kalenjin, and more.
## Status
| Language | Entries | Status |
|----------|---------|--------|
| Dholuo | building | 🟢 Active |
| Kikuyu | — | 🔜 Planned |
| Kamba | — | 🔜 Planned |
## Licences
- **Code:** MIT
- **Data:** CC BY 4.0
## Contributing
Read CONTRIBUTING.md to get started. No linguistics degree required — if you speak, you qualify.
## Community
Questions, ideas, and discussions live in GitHub Discussions.