Logo Lanfrica

skevin-dev/Data-Engineering-Speech-to-text-data-collection

Domaine:

natural language processing

Type de record:

project
Créateur:
ske
Hôte:
The purpose of this project is to build a data engineering pipeline that allows recording millions of Amharic and Swahili speakers reading digital texts in-app and web platforms. African language Speech Recognition - Speech-to-Text **Table of content** - Introduction - Project Structure - Installation - Author ## Introduction The purpose of this project is to build a data engineering pipeline that allows recording millions of Amharic and Swahili speakers reading digital texts in-app and web platforms. For this project, the Amharic news text classification dataset with baseline performance dataset is used. The aim of this project is to produce a tool that can be deployed to process posting and receiving text and audio files from and into a data lake, apply transformation in a distributed manner, and load it into a warehouse in a suitable format to train a speech-to-text model. ## Project Structure There are several files in the repository, including Python scripts, Jupyter notebooks,  and text files.  ## Installation ``` git clone github.com Cd Data-Engineering-Speech-to-text-data-collection Jupyter notebook ``` ## Author 👤 **Shyaka Kevin** - GitHub: Shyaka Kevin - LinkedIn: Shyaka Kevin