Logo Lanfrica

Zeinab-Haroon/Speech-Recognition-Project

Domain:

natural language processing

Record type:

dataset
Creator:
Zei
Host:
Speech recognition Project for low resource language(Arabic) # Speech-Recognition-Project in Arabic Language ## Taught by Gabriel, Neil, Emmanuel and Laurent Besacier. The project aims to build small Automatic Speech Recognition in high resource language (Arabic). Linear layers were trained on top of pre-trained contrastive predictive coding (CPC) and it was fine-tuned with connection temporal classification (CTC). As a result, the Character error rate (CER) and the Phone error rate (PER) for both the train and test Dataset have been computed. ## Data Collection: The Data contains two hours of recording sets that were collected using the LIG-Aikuma application in Arabic Language. The Data has been split into train (wav files)- 40 mins, test- one hour and validation set in 20 mins. The recording has been done by Zeinab Haroon in different resources from Culture studies, Sudanese novels, short stories, feminism and various articles.