Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Multimedia Digital Resources for African Tone Languages

Domain:

natural language processing

Record type:

dataset
Creator:
OloOla
Editor:
Oba
Publisher:
Men
Host:avatar
The dataset is a multimodal collection comprising digital images and natural-language speech recordings associated with selected food and agricultural product categories. The dataset was developed to support research and applications in areas such as multimedia computing, computer vision, speech processing, natural language processing, multilingual computing, image classification, speech recognition, and indigenous language technology. The dataset is organized into two principal folders: Image and Sound. The Image folder contains categorized visual representations of food and agricultural products, while the Sound folder contains spoken-word pronunciations corresponding to the represented items in four languages: English, Yoruba, Hausa, and Igbo. The Image folder contains a total of 94 image files in JPG format, systematically distributed across six food and agricultural product categories: cereal, fruit, grain, herb, tuber, and vegetable. The cereal category contains 8 images, the fruit category contains 26 images, the grain category contains 12 images, the herb category contains 18 images, while the tuber and vegetable categories contain 15 images each. Thus, the six categories collectively constitute the 94 JPG image files in the dataset. Each image represents an item belonging to one of the specified food and agricultural product categories. The category-based organization provides a structured visual dataset that can facilitate research and experimentation in image recognition, visual object retrieval, computer vision, and multimodal learning tasks. The Sound folder contains 175 natural-language speech recordings corresponding to the food and agricultural products represented in the image collection. The speech recordings provide verbal pronunciations of the associated items in four languages: English, Yoruba, Hausa and Igbo. All sound recordings are stored in WAV (.wav) format. The multilingual structure makes the dataset potentially useful for research involving multilingual speech recognition, pronunciation modelling, speech-to-text systems, African language technologies, language identification, and image–speech association. The dataset is arranged hierarchically to make individual categories and languages easily identifiable. Users can access the Image folder to obtain the visual component and the Sound folder to obtain the multilingual speech component. The Sound folder contains a total of 175 natural-language speech recordings associated with the food and agricultural products represented in the Image folder. The Sound folder is organized first according to the six food and agricultural product categories: cereal, fruit, grain, herb, tuber, and vegetable. Within each category folder, four language-specific subfolders are provided: English, Yoruba, Hausa, and Igbo. These language folders contain spoken-word pronunciations of the corresponding items in the respective languages. All speech recordings are stored in .wav format.

Visit

doi.org

Tasks

image classificationspeech processingcomputer vision

Languages

HausaIgboYoruba

Tags

Computer VisionComputer in EducationDigital EducationMultimodal Language Model

Licenses

info:eu-repo/semantics/openAccessCreative Commons Attribution 4.0 Internationalhttps://creativecommons.org/licenses/by/4.0/legalcode