Contains swahili speech recordings
# Swahili Speech Recognition Data
This repository contains speech data collected using Lig-Aikuma. It contains approximately 2hrs of Swahili speech togther with their correspoding text.
# Downloading
To clone this repo go to your terminal and type:
```
git clone
github.com
```
This adds the folder Swahili-Speech to your local directory
# Data Collection
The data collected for this project was collected from the VOASwahili website. The text used for the recordings had approximately 500 sentences with about 20000+ words/tokens.
# Data
In the data sub-folder we have:
- chars.txt
- charset.txt
- linker.txt
- raw_text_file.txt
- records(contains folders with train, test and validation)
# License
MIT