Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

thiomi_5k_v1

Domain:

natural language processing

Record type:

dataset
Creator:
thi
Host:
This dataset release is part of the data and modeling effort described in: The Thiomi Dataset: A Large-Scale Multimodal Corpus for Low-Resource African Languages The broader Thiomi work focuses on building high-quality multimodal resources for African languages and enabling ASR/MT/TTS research and deployment. Per language config (eng_Latn, kik_Latn, kam_Latn, luo_Latn, mer_Latn, som_Latn):

Visit

huggingface.co

Tasks

speech processing

Languages

DholuoGikuyuKambaKimîîruSomali

Tags

speechaudiottsasrmultilinguallow-resourceafrican-languages