Abstract
Automatic dialect identification is important for developing inclusive speech technologies in low-resource languages. Somali includes several major dialect varieties, yet dialect-level speech classification remains underexplored. This study presents a preliminary deep learning approach for classifying Maxaatiri/Standard Somali and Maay Somali speech using mel-spectrogram acoustic features and a convolutional neural network (CNN). A custom audio collection pipeline was developed to extract speech segments from publicly accessible Somali-language YouTube broadcast videos, with Maxaatiri/Standard Somali samples collected from Somali National Television and Maay Somali samples collected from Arlaadi TV. The audio was converted to mono-channel 16 kHz WAV format, processed through silence-based segmentation, divided into fixed-length speech chunks, and transformed into mel-spectrogram representations for model training. The final experimental dataset contained 180 audio segments, balanced across the two dialect classes. The CNN achieved an accuracy of 81.48%, precision of 81.60%, recall of 81.48%, and weighted F1-score of 81.43% on a held-out test set. The confusion matrix showed that the model correctly classified most samples from both dialect categories, although cross-dialect misclassification remained. These findings provide an initial computational baseline for Somali dialect detection and demonstrate the feasibility of deep learning-based acoustic classification for Somali speech. However, because the current dataset was derived from limited broadcast sources, contained a small number of samples, and lacked speaker-level metadata, the reported performance should be interpreted as a preliminary baseline rather than evidence of generalization across all Somali speakers, dialect communities, or recording conditions.