The dataset comprises three components: audio clips, an audio mapping file, and raw audio of Bomitaba, a Bantu language spoken in the Congo. Each audio clip is paired with its corresponding transcription. There are 2,613 transcribed audio clips, totalling 182 minutes and 4 seconds. There are two raw audio files totalling 121 minutes and 14.24 seconds. The audio mapping file contains 2,610 lines. Each line begins with the name of an audio file, followed by a tab, then the corresponding text exce