Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

KinyaColBERT: A Lexically Grounded Retrieval Model for Low-Resource Retrieval-Augmented Generation

Domain:

natural language processing

Record type:

papermodelsoftware
Creator:
NzeRub
Host:avatar
The recent mainstream adoption of large language model (LLM) technology is enabling novel applications in the form of chatbots and virtual assistants across many domains. With the aim of grounding LLMs in trusted domains and avoiding the problem of hallucinations, retrieval-augmented generation (RAG) has emerged as a viable solution. In order to deploy sustainable RAG systems in low-resource settings, achieving high retrieval accuracy is not only a usability requirement but also a cost-saving strategy. Through empirical evaluations on a Kinyarwanda-language dataset, we find that the most limiting factors in achieving high retrieval accuracy are limited language coverage and inadequate sub-word tokenization in pre-trained language models. We propose a new retriever model, KinyaColBERT, which integrates two key concepts: late word-level interactions between queries and documents, and a morphology-based tokenization coupled with two-tier transformer encoding. This methodology results in lexically grounded contextual embeddings that are both fine-grained and self-contained. Our evaluation results indicate that KinyaColBERT outperforms strong baselines and leading commercial text embedding APIs on a Kinyarwanda agricultural retrieval benchmark. By adopting this retrieval strategy, we believe that practitioners in other low-resource settings can not only achieve reliable RAG systems but also deploy solutions that are more cost-effective.

Visit

arxiv.org

Tasks

information retrieval

Languages

Kinyarwanda

Tags

Computation and LanguageI.2.7; I.2

Similar

BengaliMCQ: Structure-Aware Retrieval-Augmented Generation for MCQ Generation and Answer Prediction in a Low-Resource LanguageCross-Lingual Retrieval Augmented Prompt for Low-Resource LanguagesCurriculum-Aware Retrieval-Augmented Generation for Bilingual Tutoring in Low-Resource Swahili–English Secondary SchoolsCost-Efficient Cross-Lingual Retrieval-Augmented Generation for Low-Resource Languages: A Case Study in Bengali Agricultural AdvisoryMultilingual Retrieval-Augmented Generation for Knowledge-Intensive TaskAn Evidence-Grounded Retrieval-Augmented Transformer Framework for Health Misinformation Verification

BengaliMCQ: Structure-Aware Retrieval-Augmented Generation for MCQ Generation and Answer Prediction in a Low-Resource Language

Traditional retrieval-augmented generation (RAG) frameworks process documents without attending to t

Cross-Lingual Retrieval Augmented Prompt for Low-Resource Languages

Multilingual Pretrained Language Models (MPLMs) have shown their strong multilinguality in recent em

Curriculum-Aware Retrieval-Augmented Generation for Bilingual Tutoring in Low-Resource Swahili–English Secondary Schools

In Tanzanian secondary education, Swahili-language-based question-answering systems currently face s

Cost-Efficient Cross-Lingual Retrieval-Augmented Generation for Low-Resource Languages: A Case Study in Bengali Agricultural Advisory

Access to reliable agricultural advisory remains limited in many developing regions due to a persist

Multilingual Retrieval-Augmented Generation for Knowledge-Intensive Task

Retrieval-augmented generation (RAG) has become a cornerstone of contemporary NLP, enhancing large l

An Evidence-Grounded Retrieval-Augmented Transformer Framework for Health Misinformation Verification

The rapid spread of false and misleading health information through digital platforms has become a m