Logo Lanfrica

Oro_Word

Domaine:

natural language processing

Type de record:

dataset
Hôte:
This dataset contains word-level recordings in Afaan Oromoo collected from native speakers to support the development of open-source speech technologies. The dataset is designed for training and evaluating automatic speech recognition (ASR) and text-to-speech (TTS) systems. Each audio file is paired with its corresponding written word and metadata. Afaan Oromoo is a widely spoken Cushitic language in Ethiopia and neighboring regions, but it remains underrepresented in digital language resources. This contribution aims to expand accessible linguistic data, support research and education, and strengthen the presence of Afaan Oromoo in modern AI technologies.