This dataset supports the experimental evaluation of a hybrid spelling correction framework designed for low-resource languages, specifically focusing on the Coastal dialect of the Rejang language. The dataset contains 1,000 test tokens subjected to keyboard proximity error simulations.
It includes comparative performance metrics evaluating the proposed character-level N-gram and Euclidean distance combination against industry-standard benchmarks, namely Levenshtein distance and Jaro-Winkler distance. The evaluation metrics focus on system accuracy, precision, recall, and F1-score to analyze model stability and phonetic variation adaptation under constrained resource scenarios.