The purpose of this project is to build a data engineering pipeline that allows recording millions of Amharic and Swahili speakers reading digital texts in-app and web platforms.
African language Speech Recognition - Speech-to-Text
**Table of content**
- Introduction
- Project Structure
- Installation
- Author
## Introduction
The purpose of this project is to build a data engineering pipeline that allows recording millions of Amharic and Swahili speakers reading digital texts in-app and web platforms. For this project, the Amharic news text classification dataset with baseline performance dataset is used.
The aim of this project is to produce a tool that can be deployed to process posting and receiving text and audio files from and into a data lake, apply transformation in a distributed manner, and load it into a warehouse in a suitable format to train a speech-to-text model.
## Project Structure
There are several files in the repository, including Python scripts, Jupyter notebooks, and text files.
## Installation
```
git clone
github.com
Cd Data-Engineering-Speech-to-text-data-collection
Jupyter notebook
```
## Author
👤 **Shyaka Kevin**
- GitHub: Shyaka Kevin
- LinkedIn: Shyaka Kevin