Yoruba Text C3 is the largest Yoruba texts collected and used to train FastText embeddings in the YorubaTwi Embedding paper: https://www.aclweb.org/anthology/2020.lrec-1.335/
This dataset contains 1647022 parallel speech-text pairs for Yoruba, a language spoken primarily in
A listing of Web resources containing Interlinear Glossed Text for the language Yoruba: 23 document(
Yorùbá language training text for NLP, ASR and TTS tasks # Yorùbá text This repository contains fu
Speech recognition, text-to-speech synthesis, voice assistants, language modeling Notes / challenge
[cnua=> candidate chatroom Studio™ community sourced, open source, production resource] [forked=>Bui