Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Testing Dataset for Rejang Language Spelling Correction Using Hybrid Euclidean Distance and N-Gram Methods

Domain:

natural language processing

Record type:

dataset
Creator:
Sas
Publisher:
Zenodo
Host:avatar
This dataset supports the experimental evaluation of a hybrid spelling correction framework designed for low-resource languages, specifically focusing on the Coastal dialect of the Rejang language. The dataset contains 1,000 test tokens subjected to keyboard proximity error simulations. It includes comparative performance metrics evaluating the proposed character-level N-gram and Euclidean distance combination against industry-standard benchmarks, namely Levenshtein distance and Jaro-Winkler distance. The evaluation metrics focus on system accuracy, precision, recall, and F1-score to analyze model stability and phonetic variation adaptation under constrained resource scenarios.

Visit

doi.orgzenodo.org

Licenses

Creative Commons Attribution 4.0 Internationalhttps://creativecommons.org/licenses/by/4.0/legalcode

Similar

Automatic Spelling Corrector for Yorùbá Language Using Edit Distance and N-Gram Language ModelsAutomatic spelling error detection and correction for Tigrigna information retrieval: a hybrid approachTekleab15/N-gram-Language-ModelsOptimizing n‑gram Order of an n‑gram Based Language Identification Algorithm for 68 Written Languagespnabende/spelling-correction-for-East-African-languagesJONAHKYAGABA/N-gram-Language-Model-Construction-for-Oromo-Amharic

Automatic Spelling Corrector for Yorùbá Language Using Edit Distance and N-Gram Language Models

Automatic spelling error detection and correction for Tigrigna information retrieval: a hybrid approach

This paper proposes a hybrid approach to design and implement query spelling error detection and cor

Tekleab15/N-gram-Language-Models

A project to create and analyze n-gram language models using Amharic corpus # N-gram Language Model

Optimizing n‑gram Order of an n‑gram Based Language Identification Algorithm for 68 Written Languages

Language identification technology is widely used in the domains of machine learning and text mining

pnabende/spelling-correction-for-East-African-languages

# spelling-correction-for-East-African-languages This repository contains synthetic word typos for t

JONAHKYAGABA/N-gram-Language-Model-Construction-for-Oromo-Amharic

N-gram Language Model Construction for Oromo & Amharic # 📚 N-gram Language Model Construction for O