Logo Lanfrica

TonyN2060/Wav2vec_swahili_model

Domaine:

natural language processing

Type de record:

model
Créateur:
Ton
Hôte:
Fine_tuning Facebook's XLSR wav2vec transcriber model for the Swahili language # Wav2vec_swahili_model Fine_tuning Facebook's XLSR wav2vec transcriber model for the Swahili language # Speech-to-Text using XLSR-Wav2Vec2 This repository contains code for training, evaluating, and performing inference with a Speech-to-Text (STT) model using the XLSR-Wav2Vec2 architecture. The model is trained on the Common Voice Swahili dataset. ## Getting Started ### Prerequisites - Python 3.6 or later - Google Colab (for notebook execution) - Transformers - Datasets - PyTorch - Accelerate - torchaudio - librosa - jiwer ## Installing Dependencies pip install transformers datasets torch jiwer accelerate sentencepiece torchaudio librosa # Dataset Preparation The Common Voice Swahili dataset is used for training and evaluation. The dataset is loaded using the Hugging Face datasets library. # Model Architecture The XLSR-Wav2Vec2 model is used for the STT task. The model is pretrained and fine-tuned on the Swahili Common Voice dataset. # Training The training process involves loading and preprocessing the audio data, augmenting the audio data for improved model generalization, tokenizing the text using a custom vocabulary, feature extraction using the XLSR-Wav2Vec2 processor, and training the model with gradient checkpointing enabled. Training parameters, such as batch size, learning rate, and evaluation strategy, can be customized in the TrainingArguments. # Evaluation The model is evaluated on a separate validation set using the Word Error Rate (WER) metric. The trained model is then saved for future use. # Inference Inference is performed on a test set, and the transcriptions are saved to a CSV file. # Results The trained model achieves competitive performance on the validation set, as measured by the WER metric. # Contributors Tony Ndung'u Munene Wayne Otwori Maranga