Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

WDYW.01: Tone Restoration for Standard Yorùbá

Domain:

natural language processing

Record type:

softwarepaper
Creator:
Ake
Publisher:
Ace
Host:avatar
Purpose: This paper addresses automatic tone restoration for Standard Yorùbá, a three-tone language in which tonal contrasts are lexically distinctive yet the overwhelming majority of digital Yorùbá text is produced without tone marks. It presents ToneMarker, the first module (WDYW.01) of the WDYW research program (Well-formed Digital Yorùbá Writing), and evaluates its accuracy on an independently constructed test set. Design/methodology/approach: ToneMarker (afrolangtech.acelearn247.com) is a freely accessible, browser-based tool built on an 11.8-million-word verified corpus, implementing a four-layer architecture (expert-verified corrections, sentence-level exact matching, a trigram-based context model, and frequency-based fallback), extended with a rule-based disambiguation layer for high-frequency monosyllabic words grounded in Awobuluyi's (2022) grammar of Standard Yorùbá. The system was evaluated against a 48-sentence gold-standard test set sourced from Yorùbá Wikipedia and tone-marked entirely by the author's own linguistic judgement, without reference to the tool's output, using automated browser interaction for fully reproducible scoring. Findings: ToneMarker achieves 77.1% word-level accuracy (95% CI: 74.1–79.8%) on the independent evaluation set, rising to 78.3% (624/797) when the four-way homophony of ni/ní, for which no syntactic or morphological resolution could be identified in the reference grammar consulted, is excluded with explicit justification. The paper also reports a corpus-engineering finding, a double-diacritic-stripping bug in standard tokenisation pipelines that silently corrupts 49.1% of word-final double-diacritic tokens in Yorùbá text, and a companion evaluation-methodology finding that non-canonical Unicode composition order in double-diacritic vowels can distort gold-standard comparisons by tens of percentage points unless explicitly normalised. Originality: To the author's knowledge, this is the first Yorùbá tone-restoration system to combine corpus-scale statistical modelling with a citable, grammar-grounded disambiguation layer, and the first published account to document, quantify, and correct the double-diacritic stripping bug affecting standard NLP tokenisation pipelines for Yorùbá. The treatment of the ni/ní homophony as a case of principled grammatical irreducibility, rather than forcing an arbitrary default, was not identified in the reviewed literature. Contribution to the field: The paper offers a freely deployable, browser-based tool requiring no installation or specialised hardware, immediately usable by Yorùbá scholars, students, and language professionals. It sets out a scoped roadmap for comprehensive digital Yorùbá writing support: lexical correction (WDYW.02) as a near-term extension and grammatical correction as a longer-term direction, while arguing that continued use of the present system is itself a step toward the verified training data a future neural approach to Yorùbá tone restoration would require. Keywords: Yorùbá NLP, Automatic Diacritic Restoration, Tone Mark, Low-resource Language, Grammar-grounded Disambiguation, AfroLangTech

Visit

doi.org

Tasks

diacritic restorationtext normalization

Languages

Yoruba

Tags

Yorùbá NLP, Automatic Diacritic Restoration, Tone Mark, Low-resource Language, Grammar-grounded Disambiguation, AfroLangTech

Licenses

Creative Commons Attribution 4.0 Internationalhttps://creativecommons.org/licenses/by/4.0/legalcodeCopyright (C) 2026 Tosin Sina Akerele.http://rightsstatements.org/vocab/InC/1.0/