Logo Lanfrica

CateGitau/Swahili-Speech

Domain:

natural language processing

Record type:

dataset
Creator:
Cat
Host:
Contains swahili speech recordings # Swahili Speech Recognition Data This repository contains speech data collected using Lig-Aikuma. It contains approximately 2hrs of Swahili speech togther with their correspoding text. # Downloading To clone this repo go to your terminal and type: ``` git clone github.com ``` This adds the folder Swahili-Speech to your local directory # Data Collection The data collected for this project was collected from the VOASwahili website. The text used for the recordings had approximately 500 sentences with about 20000+ words/tokens. # Data In the data sub-folder we have: - chars.txt - charset.txt - linker.txt - raw_text_file.txt - records(contains folders with train, test and validation) # License MIT