This Tumbuka Language dataset, mainly focuses on datasets that are to be used for fine-tuning already existing AI Models, so that they are able to understand the Tumbuka Bantu Language (Malawi & Zambia & Tanzania).
The datasets are in different formats, and sometimes you will notice that the same dataset, have been uploaded with several file formats like .json and .csv, you just need to pick the one that suites your needs.