Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

WDYW.01: Tone Restoration for Standard Yorùbá

Domaine:

natural language processing

Type de record:

softwarepaper
Créateur:
Ake
Éditeur:
Ace
Hôte:avatar
Purpose: This paper addresses automatic tone restoration for Standard Yorùbá, a three-tone language in which tonal contrasts are lexically distinctive yet the overwhelming majority of digital Yorùbá text is produced without tone marks. It presents ToneMarker, the first module (WDYW.01) of the WDYW research program (Well-formed Digital Yorùbá Writing), and evaluates its accuracy on an independently constructed test set. Design/methodology/approach: ToneMarker (afrolangtech.acelearn247.com) is a freely accessible, browser-based tool built on an 11.8-million-word verified corpus, implementing a four-layer architecture (expert-verified corrections, sentence-level exact matching, a trigram-based context model, and frequency-based fallback), extended with a rule-based disambiguation layer for high-frequency monosyllabic words grounded in Awobuluyi's (2022) grammar of Standard Yorùbá. The system was evaluated against a 48-sentence gold-standard test set sourced from Yorùbá Wikipedia and tone-marked entirely by the author's own linguistic judgement, without reference to the tool's output, using automated browser interaction for fully reproducible scoring. Findings: ToneMarker achieves 77.1% word-level accuracy (95% CI: 74.1–79.8%) on the independent evaluation set, rising to 78.3% (624/797) when the four-way homophony of ni/ní, for which no syntactic or morphological resolution could be identified in the reference grammar consulted, is excluded with explicit justification. The paper also reports a corpus-engineering finding, a double-diacritic-stripping bug in standard tokenisation pipelines that silently corrupts 49.1% of word-final double-diacritic tokens in Yorùbá text, and a companion evaluation-methodology finding that non-canonical Unicode composition order in double-diacritic vowels can distort gold-standard comparisons by tens of percentage points unless explicitly normalised. Originality: To the author's knowledge, this is the first Yorùbá tone-restoration system to combine corpus-scale statistical modelling with a citable, grammar-grounded disambiguation layer, and the first published account to document, quantify, and correct the double-diacritic stripping bug affecting standard NLP tokenisation pipelines for Yorùbá. The treatment of the ni/ní homophony as a case of principled grammatical irreducibility, rather than forcing an arbitrary default, was not identified in the reviewed literature. Contribution to the field: The paper offers a freely deployable, browser-based tool requiring no installation or specialised hardware, immediately usable by Yorùbá scholars, students, and language professionals. It sets out a scoped roadmap for comprehensive digital Yorùbá writing support: lexical correction (WDYW.02) as a near-term extension and grammatical correction as a longer-term direction, while arguing that continued use of the present system is itself a step toward the verified training data a future neural approach to Yorùbá tone restoration would require. Keywords: Yorùbá NLP, Automatic Diacritic Restoration, Tone Mark, Low-resource Language, Grammar-grounded Disambiguation, AfroLangTech

Visit

doi.org

Tasks

diacritic restorationtext normalization

Languages

Yoruba

Tags

Yorùbá NLP, Automatic Diacritic Restoration, Tone Mark, Low-resource Language, Grammar-grounded Disambiguation, AfroLangTech

Licenses

Creative Commons Attribution 4.0 Internationalhttps://creativecommons.org/licenses/by/4.0/legalcodeCopyright (C) 2026 Tosin Sina Akerele.http://rightsstatements.org/vocab/InC/1.0/

Similaires

Improving Yorùbá Diacritic RestorationAutomatic Diacritic Restoration of Yorùbá language TextAttentive Sequence-to-Sequence Learning for Diacritic Restoration of Yorùbá Language TextYoruba-G2P: A tone-aware grapheme-to-phoneme converter for YorùbáADDRESSING TONE-MARKING CHALLENGES IN DIGITISED YORÙBÁ TEXTS FOR WEB COMMUNICATIONSMichka Sachnine, Grammaire du yorùbá standard (Nigéria), 2014

Improving Yorùbá Diacritic Restoration

Yorùbá is a widely spoken West African language with a writing system rich in orthographic and tonal diacritics. They provide morphological information, are crucial for lexical disambiguation, pronunciation and are vital for any computational Speech or Natural Lang

Automatic Diacritic Restoration of Yorùbá language Text

Automatic Diacritic Restoration of Yorùbá language Text

Attentive Sequence-to-Sequence Learning for Diacritic Restoration of Yorùbá Language Text

Yorùbá is a widely spoken West African language with a writing system rich in tonal and orthographic diacritics. With very few exceptions, diacritics are omitted from electronic texts, due to limited device and application support. Diacritics provide morphological

Yoruba-G2P: A tone-aware grapheme-to-phoneme converter for Yorùbá

Yoruba-G2P is a tone-aware grapheme-to-phoneme converter for Yorùbá, designed to support speech and

ADDRESSING TONE-MARKING CHALLENGES IN DIGITISED YORÙBÁ TEXTS FOR WEB COMMUNICATIONS

Abstract This study is driven by the following questions: Do Yoruba born-digital datasets consiste

Michka Sachnine, Grammaire du yorùbá standard (Nigéria), 2014

International audience