This is dataset of speech translation task for Bemba-to-English Language. This dataset is acquired from (Big-C)[BIG-C: Bemba Image Grounded…] github repository.
Big-C is a large conversations dataset between Bemba Speakers based on Image [1]. This dataset provide data for speech translation.
Some preprocessing was done in this dataset.
Drop some unused columns other than audio_id, sentence, translation, and speaker_id.