For this project entitled “Development of a speech recognition model for under-resourced languages”, the collected dataset includes audio recordings of native speakers of the Dioula language pronouncing a series of specific keywords such as “Yes”, “No”, “Call”, “Next”, “Previous”, “Dial”, “Stop”, “True”, “False”, “Enter” and “Exit”. These recordings were obtained in diverse environments, with a wide variety of speakers to ensure that the different tones, accents, and phonetic variations specific to Dioula are taken into account. Each audio file was assigned to a category corresponding to the spoken word, creating a class-structured dataset. This corpus will be used to train a speech recognition model using deep learning techniques, with the objective of improving recognition accuracy in real contexts. This database is essential for the design and evaluation of a model capable of facilitating access and voice interaction for the Dioula community.