Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Automatic Correction of Writing Anomalies in Hausa Texts

Domaine:

natural language processing

Type de record:

paperdatasetmodel
Créateur:
WalNis
Hôte:avatar
Hausa texts are often characterized by writing anomalies, such as incorrect character substitutions and spacing errors, which sometimes hinder natural language processing (NLP) applications. This paper presents an approach to automatically correct anomalies by finetuning transformer-based models. Using a corpus gathered from several public sources, we create a large-scale parallel dataset of over 400,000 noisy-clean Hausa sentence pairs by introducing synthetically generated noise to mimic realistic writing errors. In addition, we finetune several multilingual and African language models, including M2M100, AfriTeVA, NCAIR1/N-ATLaS, UBC-NLP/cheetah-base, and other variants of BART and T5 for this correction task. Our experimental results demonstrate that models such as M2M100 achieve state-of-the-art results despite their smaller size and distinct pretraining, and that correcting errors can have a significant impact in improving downstream tasks such as text classification, machine translation, question answering, and LLM prompting in general. This research provides a methodology, a publicly available dataset, and a comparison of models to improve Hausa text quality, thereby advancing NLP capabilities for the language and offering transferable insights for other low-resource languages. Accepted at ACL2026

Visit

arxiv.org

Tasks

grammar error correction

Languages

Hausa

Tags

Computation and Language

Similaires

ahmadmwali/hausa-writing-anomaliesAutomatic authorship attribution in Albanian textsAutomatic Correction of Indonesian Grammatical Errors Based on TransformerOCR Post Correction for Endangered Language TextsAutomatic generator of mathematical logic questions in French with automated correction

ahmadmwali/hausa-writing-anomalies

Automatic authorship attribution in Albanian texts

Automatic authorship identification is a challenging task that has been the focus of extensive resea

Automatic Correction of Indonesian Grammatical Errors Based on Transformer

Grammatical error correction (GEC) is one of the major tasks in natural language processing (NLP) wh

OCR Post Correction for Endangered Language Texts

There is little to no data available to build natural language processing models for most endangered languages. However, textual data in these languages often exists in formats that are not machine-readable, such as paper books and scanned images. In this work, we

Automatic generator of mathematical logic questions in French with automated correction

Automating the generation of questions to assess students remains a serious scientific problem in de