The dataset we have compiled for our research on "Setting up a speech recognition model for under-resourced languages" consists of audio recordings of Dioula speakers pronouncing the numbers 1, 2, 3, and 4. These recordings were collected under various conditions, featuring variability in speakers, accents, and environmental contexts. The data has been categorized into four distinct classes, each corresponding to one of the numbers (1, 2, 3, or 4), enabling the training and evaluation of a machine learning-based speech recognition model.