The Common Voice dataset consists of a unique MP3 and corresponding text file. Many of the 30328 rec
English-Hausa parallel corpus.