Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Yorùbá Text C3

Domain:

natural language processing

Record type:

dataset
Creator:
aje
Host:
Yoruba Text C3 is the largest Yoruba texts collected and used to train FastText embeddings in the YorubaTwi Embedding paper: aclweb.org

Visit

huggingface.co

Tasks

embeddings

Languages

Yoruba

Licenses

cc-by-nc-4.0

Similar

Yoruba Twi Text C3Automatic Diacritic Restoration of Yorùbá language TextText Detoxification in isiXhosa and Yorùbá DatasetMachine Translation System for Numeral in English Text to Yorùbá LanguageAttentive Sequence-to-Sequence Learning for Diacritic Restoration of Yorùbá Language Texthresholding in Convolutional Neural Network Model for Yorùbá Speech-to-Text Conversion

Yoruba Twi Text C3

Yoruba Text C3 is the largest Yoruba texts collected and used to train FastText embeddings in the YorubaTwi Embedding paper: https://www.aclweb.org/anthology/2020.lrec-1.335/

Automatic Diacritic Restoration of Yorùbá language Text

Automatic Diacritic Restoration of Yorùbá language Text

Text Detoxification in isiXhosa and Yorùbá Dataset

A parallel dataset of toxic and detoxified sentence pairs (178 pairs each) was manually generated fo

Machine Translation System for Numeral in English Text to Yorùbá Language

The machine translation of numbers from the English language into the Yorùbá language is an integral

Attentive Sequence-to-Sequence Learning for Diacritic Restoration of Yorùbá Language Text

Yorùbá is a widely spoken West African language with a writing system rich in tonal and orthographic diacritics. With very few exceptions, diacritics are omitted from electronic texts, due to limited device and application support. Diacritics provide morphological

hresholding in Convolutional Neural Network Model for Yorùbá Speech-to-Text Conversion

Speech-to-text and text-to-speech conversions are referred to as Automatic Speech Recognition (ASR),