Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Leveraging Retrieval-Augmented Generation for Swahili Language Conversation Systems

Domaine:

natural language processing

Type de record:

paper
Créateur:
EdmQinGimXu,
Éditeur:
MDP
Hôte:
A conversational system is an artificial intelligence application designed to interact with users in natural language, providing accurate and contextually relevant responses. Building such systems for low-resource languages like Swahili presents significant challenges due to the limited availability of large-scale training datasets. This paper proposes a Retrieval-Augmented Generation-based system to address these challenges and improve the quality of Swahili conversational AI. The system leverages fine-tuning, where models are trained on available Swahili data, combined with external knowledge retrieval to enhance response accuracy and fluency. Four models—mT5, GPT-2, mBART, and GPT-Neo—were evaluated using metrics such as BLEU, METEOR, Query Performance, and inference time. Results show that Retrieval-Augmented Generation consistently outperforms fine-tuning alone, particularly in generating detailed and contextually appropriate responses. Among the tested models, mT5 with Retrieval-Augmented Generation demonstrated the best performance, achieving a BLEU score of 56.88%, a METEOR score of 72.72%, and a Query Performance score of 84.34%, while maintaining relevance and fluency. Although Retrieval-Augmented Generation introduces slightly longer response times, its ability to significantly improve response quality makes it an effective approach for Swahili conversational systems. This study highlights the potential of Retrieval-Augmented Generation to advance conversational AI for Swahili and other low-resource languages, with future work focusing on optimizing efficiency and exploring multilingual applications.

Visit

doi.org

Languages

Swahili

Licenses

https://creativecommons.org/licenses/by/4.0/

Similaires

Multilingual Retrieval-Augmented Generation for Knowledge-Intensive TaskCurriculum-Aware Retrieval-Augmented Generation for Bilingual Tutoring in Low-Resource Swahili–English Secondary SchoolsKinyaColBERT: A Lexically Grounded Retrieval Model for Low-Resource Retrieval-Augmented GenerationBengaliMCQ: Structure-Aware Retrieval-Augmented Generation for MCQ Generation and Answer Prediction in a Low-Resource LanguageLeveraging AI chat assistants for enhanced food security in Africa: A comprehensive integration of large language models, retrieval augmented generation, and vector embedding techniquesError-Aware TF-IDF Retrieval-Augmented Generation for ASR Error Correction

Multilingual Retrieval-Augmented Generation for Knowledge-Intensive Task

Retrieval-augmented generation (RAG) has become a cornerstone of contemporary NLP, enhancing large l

Curriculum-Aware Retrieval-Augmented Generation for Bilingual Tutoring in Low-Resource Swahili–English Secondary Schools

In Tanzanian secondary education, Swahili-language-based question-answering systems currently face s

KinyaColBERT: A Lexically Grounded Retrieval Model for Low-Resource Retrieval-Augmented Generation

The recent mainstream adoption of large language model (LLM) technology is enabling novel applicatio

BengaliMCQ: Structure-Aware Retrieval-Augmented Generation for MCQ Generation and Answer Prediction in a Low-Resource Language

Traditional retrieval-augmented generation (RAG) frameworks process documents without attending to t

Leveraging AI chat assistants for enhanced food security in Africa: A comprehensive integration of large language models, retrieval augmented generation, and vector embedding techniques

Agricultural productivity in Africa faces significant challenges due to limited access to timely, lo

Error-Aware TF-IDF Retrieval-Augmented Generation for ASR Error Correction

End-to-end automatic speech recognition systems frequently hallucinate rare entities and domain-spec