Logo Lanfrica

waleghwa/low-resource-language-data

Domain:

natural language processing

Record type:

dataset
Creator:
wal
Host:
**About** This repository is the result of work funded by the Lacuna Fund. We collected parallel corpora for Kiswahili and three indigenous Kenyan languages: Kidaw'ida, Kalenjin and Dholuo. We also uploaded the indigenous language text data to Mozilla Common Voice and crowd-source voice data by having volunteers read the text data. **How to Cite** This repository is on Zenodo and should be cited as follows: Mbogho, A., Kipkebut, A., Wanzare, L., Awuor, Q., Oloo, V., & Lugano, R. (2024). waleghwa/low-resource-language-data: Parallel Corpora for Kiswahili and Kidaw'ida, Kalenjin and Dholuo (v1.0.0) [Data set]. Zenodo. waleghwa/low-resource-langu…