Building Text and Speech Datasets for Low Resourced Languages: A Case of Languages in East Africa
Africa has over 2000 languages; however, those languages are not well represented in the existing Natural Language Processing ecosystem. African languages lack essential digital resources to be engaged effectively in the advancing language technologies. This growin