Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Automatic Identification of Maxaatiri and Maay Somali Dialects from Speech Using Mel-Spectrogram-Based Convolutional Neural Networks

Domain:

natural language processing

Record type:

datasetpaper
Creator:
Moh
Publisher:
Spr
Host:
Abstract Automatic dialect identification is important for developing inclusive speech technologies in low-resource languages. Somali includes several major dialect varieties, yet dialect-level speech classification remains underexplored. This study presents a preliminary deep learning approach for classifying Maxaatiri/Standard Somali and Maay Somali speech using mel-spectrogram acoustic features and a convolutional neural network (CNN). A custom audio collection pipeline was developed to extract speech segments from publicly accessible Somali-language YouTube broadcast videos, with Maxaatiri/Standard Somali samples collected from Somali National Television and Maay Somali samples collected from Arlaadi TV. The audio was converted to mono-channel 16 kHz WAV format, processed through silence-based segmentation, divided into fixed-length speech chunks, and transformed into mel-spectrogram representations for model training. The final experimental dataset contained 180 audio segments, balanced across the two dialect classes. The CNN achieved an accuracy of 81.48%, precision of 81.60%, recall of 81.48%, and weighted F1-score of 81.43% on a held-out test set. The confusion matrix showed that the model correctly classified most samples from both dialect categories, although cross-dialect misclassification remained. These findings provide an initial computational baseline for Somali dialect detection and demonstrate the feasibility of deep learning-based acoustic classification for Somali speech. However, because the current dataset was derived from limited broadcast sources, contained a small number of samples, and lacked speaker-level metadata, the reported performance should be interpreted as a preliminary baseline rather than evidence of generalization across all Somali speakers, dialect communities, or recording conditions.

Visit

doi.org

Tasks

language identificationspeech processing

Languages

MaaySomali

Licenses

https://creativecommons.org/licenses/by/4.0/

Similar

A Spectrogram and Local Feature-Assisted Convolutional Neural Network for Amharic Speech Emotion IdentificationAutomated speech-based screening of depression using deep convolutional neural networksAutomatic detection of indris’ songs using convolutional neural networksAutomatic Detection and Taxonomic Identification of Dolphin Vocalisations using Convolutional Neural Networks for Passive Acoustic MonitoringLanguage Identification Using Deep Convolutional Recurrent Neural NetworksAmazigh CNN speech recognition system based on Mel spectrogram feature extraction method

A Spectrogram and Local Feature-Assisted Convolutional Neural Network for Amharic Speech Emotion Identification

Abstract Speech Emotion Recognition (SER) plays a significant role in improving hu

Automated speech-based screening of depression using deep convolutional neural networks

Early detection and treatment of depression is essential in promoting remission, preventing relapse,

Automatic detection of indris’ songs using convolutional neural networks

Automatic Detection and Taxonomic Identification of Dolphin Vocalisations using Convolutional Neural Networks for Passive Acoustic Monitoring

A novel framework for evaluating dolphin sound detection and species identification is prop

Language Identification Using Deep Convolutional Recurrent Neural Networks

Language Identification (LID) systems are used to classify the spoken language from a given audio sa

Amazigh CNN speech recognition system based on Mel spectrogram feature extraction method