Logo Lanfrica

Jayem-11/Swahili_speech_to_text

Domain:

natural language processing

Record type:

model
Creator:
Jay
Host:
Speech to Text for Swahili Language with Whisper-small. # Swahili Speech to text Finetuning whisper-small for swahili speech to text. photo credits: Devonyu ## Description: This project uses the small version of Whisper, a general-purpose speech recognition model created by OpenAI, to convert Swahili audio to text. Whisper is pretrained for ASR (Automatic speech Recognition) and speech translation on 680k hours of labelled data ## Author - Github @JM_Rono - Linked_in @John Michael Rono ## Table of Contents A Data B Machine learning C Deploying ## Design ## A. Data The data consist of about 82K instances of swahili audio form Mozilla common voice. I got the dataset from participating in a Zindi competition. Join the completed competiotion to get acces to the data. ## B. Machine Learning - Check-out notebook: @notebook ### Evaluation The model had a WER score of 8.365 wandb ## C. Deploying - Deployed at: swahilispeechtotext-htxcv6s… - Upload a 30 second audio or less ## - Hit the text button.