This is a dataset project which includes a specific language in Malawi called Chitumbuka mostly used by people living in the Northen region of the country. The dataset includes the records of individuals as well as their corresponding transcriptions for TTS , training AI models and improving technological advancement in future times.