Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Enhancing Text Search Accuracy in Low-Resource Languages Using N-gram and Fuzzy Matching: A Case Study on Dari

Domaine:

natural language processing

Type de record:

paper
Créateur:
AhmSayKhoAbd
Éditeur:
LPP
Hôte:
Text search systems play an important role in modern information retrieval by providing fast and efficient access to digital information. However, despite the rapid development of digital technologies, these systems often encounter challenges when processing incomplete, misspelled, or orthographically inconsistent queries, particularly in low-resource languages such as Dari. The existence of spelling variations and limited linguistic resources reduces retrieval accuracy and negatively affects search performance. Information Retrieval (IR), as an important branch of Natural Language Processing (NLP), aims to retrieve accurate information from large text collections. Nevertheless, IR systems still face difficulties in low-resource languages due to limited datasets, insufficient computational resources, and the lack of language-specific retrieval tools. Therefore, improving search accuracy in the Dari language remains an important research challenge. This study aims to enhance text retrieval performance using the combination of n-gram analysis and Fuzzy Matching techniques. To achieve this objective, a mini search system was designed to process user queries and suggest the closest matching word whenever incomplete or misspelled inputs are detected. A Dari text dataset was collected and prepared through preprocessing, tokenization, and unique-word extraction. Fuzzy Matching was applied to measure similarity between the user query and dataset words, while n-gram analysis was used to examine structural similarity between words. The experimental findings demonstrated that the proposed mini search system successfully identified several incomplete and misspelled Dari words with relatively high similarity scores. These findings indicate that combining n-gram analysis with Fuzzy Matching can improve typo-tolerant text retrieval in low-resource languages such as Dari and provide a practical foundation for future research on Dari information retrieval systems.

Visit

doi.org

Tasks

information retrieval

Languages

Pévé

Licenses

https://creativecommons.org/licenses/by-sa/4.0

Similaires

Enhancing Conversational AI for Low-Resource Languages: A Case Study on SomaliEnhancing Pos Tagging For Low-Resource Languages: A Case Study On DholuoGemDetox at TextDetox CLEF 2025: Enhancing a Massively Multilingual Model for Text Detoxification on Low-resource LanguagesText Normalization for Low Resource LanguagesEnhancing Multilingual Table-to-Text Generation with QA Blueprints: Overcoming Challenges in Low-Resource Languages Text-to-Speech Synthesis Using Found Data for Low-Resource Languages

Enhancing Conversational AI for Low-Resource Languages: A Case Study on Somali

Conversational AI has made huge strides in understanding and generating human language. However, the

Enhancing Pos Tagging For Low-Resource Languages: A Case Study On Dholuo

GemDetox at TextDetox CLEF 2025: Enhancing a Massively Multilingual Model for Text Detoxification on Low-resource Languages

As social-media platforms emerge and evolve faster than the regulations meant to oversee them, autom

Text Normalization for Low Resource Languages

This repository contains code related to the Google open source internship project Text Normalization for Low Resource Languages.

Enhancing Multilingual Table-to-Text Generation with QA Blueprints: Overcoming Challenges in Low-Resource Languages

Limiting training data in low-resource languages is a barrier to Natural Language Processing (NLP).

Text-to-Speech Synthesis Using Found Data for Low-Resource Languages

Text-to-speech synthesis is a key component of interactive, speech-based systems. Typically, buildi