Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Somali Dialect Identification: A Low-Resource Benchmark for MAXAA TIRI and MAAY Using Machine and Deep Learning

Domaine:

natural language processing

Type de record:

dataset
Créateur:
AbdYusShaYus
Éditeur:
Spr
Hôte:
Abstract This study addresses the task of automatic dialect identification within the Somali language, focusing on its two primary dialects: MAXAA TIRI and MAAY. Somali, spoken by over 22 million individuals, presents significant dialectal diversity, which poses challenges for downstream NLP applications such as sentiment analysis, machine translation, and information retrieval. To bridge this gap, the study constructs and annotates a representative dataset of 3,011 text samples collected from diverse sources including social media and formal documents. The study evaluates a range of machine learning and deep learning models namely Naive Bayes, Support Vector Machines (SVM), and Bidirectional Long Short-Term Memory (BiLSTM)—to classify text based on dialectal features. Our results demonstrate high performance, with Naive Bayes and BiLSTM models achieving classification accuracies of 99.86% and 99.57%, respectively. To ensure generalizability, we apply rigorous validation methods, including cross-source testing and real-world deployment through a web-based interface. This research contributes a novel dataset, benchmarks several AI models for Somali dialect identification, and provides foundational insights for advancing low-resource language processing.

Visit

doi.org

Tasks

text classification

Languages

MaaySomali

Licenses

https://creativecommons.org/licenses/by/4.0/

Similaires

Somali dialect identification in low-resource settings using machine learning and deep learningWollo Dialect Identification for Amharic Language Using Machine LearningLanguage Identification in Low-Resource Multilingual Document Images using Deep Learning TechniquesVector Vigilantes: A Deep Learning Surveillance App For Mosquito Identification and Control for Low-Resource CountriesAutomatic identification of Algiers dialect based on word‑level machine learning and deep learning methodsA Deep Learning Approach for the Romanized Tunisian Dialect Identification

Somali dialect identification in low-resource settings using machine learning and deep learning

Abstract This study investigates automatic dialect identification for the Somali

Wollo Dialect Identification for Amharic Language Using Machine Learning

This thesis presents a study on Wollo dialect identification using machine learning, specifically th

Language Identification in Low-Resource Multilingual Document Images using Deep Learning Techniques

With the increasing digitization of printed materials, it has become common to encounter documents i

Vector Vigilantes: A Deep Learning Surveillance App For Mosquito Identification and Control for Low-Resource Countries

Automatic identification of Algiers dialect based on word‑level machine learning and deep learning methods

A Deep Learning Approach for the Romanized Tunisian Dialect Identification

Language identification is an important task in natural language processing that consists in determi