Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Statistical Text Analysis for Yorùbá Speech Generation Using Zipf's Law. Ife Journal of Technology

Domaine:

natural language processing

Type de record:

paper
Créateur:
Iya
Éditeur:
Fac
Hôte:avatar
The practical challenge of creating a Yorùbá text-to-speech synthesis has initiated our work on statistical text analysis. Language and speech technology applications have gained an increasingly wide-spread use in several languages/countries, and this has necessitated the importance of examining how much difference exists between English (in most cases the first language for most technologies and applications) and tone languages, specifically, Yoruba. These differences are studied and described in detail in linguistics but they rarely quantified and used by technology developers. In this paper, Yoruba language was described using text corpora from textbooks and newspapers. Other texts from Internet sources were also used. The corpus size was 291,392 word forms and the data was analyzed using Zipf. Based on the statistical analysis, it was found that the coverage of corpora by the most frequent words follows a parallel logarithmic rule for all languages in coverage range, known as Zipf’s law in linguistics.

Visit

doi.orgzenodo.org

Tasks

speech processingtext to speech

Languages

Yoruba

Tags

Text corporaData analysisStandard YorubaText-to-speech synthesis

Licenses

Creative Commons Attribution 4.0 Internationalhttps://creativecommons.org/licenses/by/4.0/legalcode

Similaires

hresholding in Convolutional Neural Network Model for Yorùbá Speech-to-Text ConversionStatistical analysis of waste generation in Bukavu city using IoT sensor dataAfrican Speech Technology Afrikaans Text CorpusAfrican Speech Technology isiZulu Text CorpusAfrican Speech Technology English Text CorpusAfrican Speech Technology isiXhosa Text Corpus

hresholding in Convolutional Neural Network Model for Yorùbá Speech-to-Text Conversion

Speech-to-text and text-to-speech conversions are referred to as Automatic Speech Recognition (ASR),

Statistical analysis of waste generation in Bukavu city using IoT sensor data

Abstract The city of Bukavu, in the Democratic Republic of Congo, faces a health a

African Speech Technology Afrikaans Text Corpus

Monolingual text corpus developed during the African Speech Technology project.

African Speech Technology isiZulu Text Corpus

Monolingual text corpus developed during the African Speech Technology project.

African Speech Technology English Text Corpus

Monolingual text corpus developed during the African Speech Technology project.

African Speech Technology isiXhosa Text Corpus

Monolingual text corpus developed during the African Speech Technology project.