Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Oro_Word

Domain:

natural language processing

Record type:

dataset
Host:
This dataset contains word-level recordings in Afaan Oromoo collected from native speakers to support the development of open-source speech technologies. The dataset is designed for training and evaluating automatic speech recognition (ASR) and text-to-speech (TTS) systems. Each audio file is paired with its corresponding written word and metadata. Afaan Oromoo is a widely spoken Cushitic language in Ethiopia and neighboring regions, but it remains underrepresented in digital language resources. This contribution aims to expand accessible linguistic data, support research and education, and strengthen the presence of Afaan Oromoo in modern AI technologies.

Visit

mozilladatacollective.com

Tasks

automatic speech recognitionspeech processingtext to speech

Languages

OromoOromo, Borana-Arsi-GujiOromo, EasternOromo, West Central

Tags

mdcmozilla data collectiveTTS.WAVCSV

Licenses

Creative Commons Zero v1.0 Universal (CC0-1.0)