First open-source Kanongesha Lunda ↔ English dataset for NLLB/Meta AI. 100 verified sentences.
# Kanongesha Lunda ↔ English Parallel Dataset v1.0
**Dialect**: Kanongesha Lunda, Northwestern Zambia
**Speakers**: ~2M native speakers, 0 representation in NLLB-200
**Size**: 100 parallel sentences, human-verified
**License**: MIT - Free for research and commercial use
**Purpose**: Enable Meta AI, NLLB, and translation tools for Lunda communities
## Why This Matters
Lunda children cannot use AI in their mother tongue. This dataset is the first step to change that.
Built for my daughter Suliya, and every Lunda speaker excluded from the digital world.
## Domains Covered
Greetings, Family, Food, Market, Health, Time, Travel, Work, Education, Weather, Home
## Data Quality
- 100% human translations by native Kanongesha Lunda speakers
- No machine translation or AI guesses
- Cultural context preserved: `nshima`, `kaloña`, `hachimu`
- Verified for NLLB-200 compatibility
## Roadmap
- [x] v1.0: 100 sentences - COMPLETE
- [ ] v1.1: 500 sentences - IN PROGRESS
- [ ] v2.0: 2000 sentences + audio - PLANNED
## Contribute
We need native speakers to add more sentences. Open a Pull Request or email me.