Fine-tuned facebook/wav2vec2-large-xlsr-53 on Fon (or Fongbe) using the Fon Dataset. When using this model, make sure that your speech input is sampled at 16kHz.