Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Grammatical Error Correction for Low-Resource Languages: The Case of Zarma

Domain:

natural language processing

Record type:

paperdataset
Creator:
KeiBreLe,Owu
Host:avatar
Grammatical error correction (GEC) aims to improve text quality and readability. Previous work on the task focused primarily on high-resource languages, while low-resource languages lack robust tools. To address this shortcoming, we present a study on GEC for Zarma, a language spoken by over five million people in West Africa. We compare three approaches: rule-based methods, machine translation (MT) models, and large language models (LLMs). We evaluated GEC models using a dataset of more than 250,000 examples, including synthetic and human-annotated data. Our results showed that the MT-based approach using M2M100 outperforms others, with a detection rate of 95.82% and a suggestion accuracy of 78.90% in automatic evaluations (AE) and an average score of 3.0 out of 5.0 in manual evaluation (ME) from native speakers for grammar and logical corrections. The rule-based method was effective for spelling errors but failed on complex context-level errors. LLMs -- Gemma 2b and MT5-small -- showed moderate performance. Our work supports use of MT models to enhance GEC in low-resource settings, and we validated these results with Bambara, another West African language.

Visit

arxiv.org

Tasks

grammar error correction

Languages

BamanankanLameZarma

Tags

Computation and LanguageMachine Learning

Similar

ChatGPT for Arabic Grammatical Error CorrectionEnhancing Text Editing for Grammatical Error Correction: Arabic as a Case StudyRobustness of Synthetic vs. Human-Annotated Grammatical Error Detection Models in Low-Resource LanguagesBeyond English: Evaluating LLMs for Arabic Grammatical Error CorrectionGrammatical Error Detection Performance in Low-Resource Languages: Zero-Shot Synthetic vs. Human-Annotated BaselinesDiversity in Zero-Shot Synthetic Data for Low-Resource Grammatical Error Detection

ChatGPT for Arabic Grammatical Error Correction

Recently, large language models (LLMs) fine-tuned to follow human instruction have exhibited signifi

Enhancing Text Editing for Grammatical Error Correction: Arabic as a Case Study

Text editing frames grammatical error correction (GEC) as a sequence tagging problem, where edit tag

Robustness of Synthetic vs. Human-Annotated Grammatical Error Detection Models in Low-Resource Languages

Grammatical Error Detection (GED) methods rely heavily on human annotated error corpora. However, th

Beyond English: Evaluating LLMs for Arabic Grammatical Error Correction

Large language models (LLMs) finetuned to follow human instruction have recently exhibited significa

Grammatical Error Detection Performance in Low-Resource Languages: Zero-Shot Synthetic vs. Human-Annotated Baselines

Grammatical Error Detection (GED) methods rely heavily on human annotated error corpora. However, th

Diversity in Zero-Shot Synthetic Data for Low-Resource Grammatical Error Detection

Grammatical Error Detection (GED) methods rely heavily on human annotated error corpora. However, th