The practical challenge of creating a Yorùbá text-to-speech synthesis has initiated our work on statistical text analysis. Language and speech technology applications have gained an increasingly wide-spread use in several languages/countries, and this has necessitated the importance of examining how much difference exists between English (in most cases the first language for most technologies and applications) and tone languages, specifically, Yoruba. These differences are studied and described in detail in linguistics but they rarely quantified and used by technology developers. In this paper, Yoruba language was described using text corpora from textbooks and newspapers. Other texts from Internet sources were also used. The corpus size was 291,392 word forms and the data was analyzed using Zipf. Based on the statistical analysis, it was found that the coverage of corpora by the most frequent words follows a parallel logarithmic rule for all languages in coverage range, known as Zipf’s law in linguistics.