Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Enhancing Conversational AI for Low-Resource Languages: A Case Study on Somali

Domaine:

natural language processing

Type de record:

paper
Créateur:
MohSer
Éditeur:
International Journal of Innovative Science and Research Technology (IJISRT)
Hôte:avatar
Conversational AI has made huge strides in understanding and generating human language. However, these advances have mostly benefited high-resource languages such as English and Spanish. In contrast, languages like Somali— spoken by an estimated 20 million people—lack the abundance of annotated data needed to develop robust language models. This study focuses on practical strategies to boost Somali text and speech processing capabilities. We explore three core approaches: (1) transfer learning, (2) synthetic data augmentation, and (3) fine-tuning multilingual models. Our experiments, featuring XLM-R, mBERT, and OpenAI’s Whisper API, show that well-adapted models significantly outperform their baseline counterparts in Somali text translation and speech-to-text tasks. Beyond the numbers, our findings underscore the societal value of creating accessible AI tools for underrepresented linguistic communities, providing a template for extending these methods to other low-resource languages.

Visit

doi.orgzenodo.org

Tasks

automatic speech recognitionmachine translationspeech processingtransfer learning

Languages

Somali

Licenses

Creative Commons Attribution 4.0 Internationalhttps://creativecommons.org/licenses/by/4.0/legalcode