Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Improving Yorùbá Diacritic Restoration

Domain:

natural language processing

Record type:

paper
Yorùbá is a widely spoken West African language with a writing system rich in orthographic and tonal diacritics. They provide morphological information, are crucial for lexical disambiguation, pronunciation and are vital for any computational Speech or Natural Language Processing tasks. However diacritic marks are commonly excluded from electronic texts due to limited device and application support as well as general education on proper usage. We report on recent efforts at dataset cultivation. By aggregating and improving disparate texts from the web and various personal libraries, we were able to significantly grow our clean Yorùbá dataset from a majority Bibilical text corpora with three sources to millions of tokens from over a dozen sources. We evaluate updated diacritic restoration models on a new, general purpose, public-domain Yorùbá evaluation dataset of modern journalistic news text, selected to be multi-purpose and reflecting contemporary usage. All pre-trained models, datasets and source-code have been released as an open-source project to advance efforts on Yorùbá language technology.

Visit

arxiv.org

Connected records

model

Tasks

diacritic restorationtext normalization

Languages

Yoruba

Tags

africanlp1

Similar

Automatic Diacritic Restoration of Yorùbá language TextAttentive Sequence-to-Sequence Learning for Diacritic Restoration of Yorùbá Language TextCrinmatic/Diacritic-RestorationNatashadonoh11/yoruba-diacritic-restorationmajtom91/yoruba-diacritic-restorationayomidedavid/yoruba-diacritic-restoration-model

Automatic Diacritic Restoration of Yorùbá language Text

Automatic Diacritic Restoration of Yorùbá language Text

Attentive Sequence-to-Sequence Learning for Diacritic Restoration of Yorùbá Language Text

Yorùbá is a widely spoken West African language with a writing system rich in tonal and orthographic diacritics. With very few exceptions, diacritics are omitted from electronic texts, due to limited device and application support. Diacritics provide morphological

Crinmatic/Diacritic-Restoration

Using AI to restore Diacritics on Yoruba language (which is a low resource language) ## **Diacritic

Natashadonoh11/yoruba-diacritic-restoration

Yoruba diacritic restoration- Adaption Labs submission # Yorùbá Diacritic Restoration Adapting a l

majtom91/yoruba-diacritic-restoration

Automatic Yorùbá diacritic restoration application # Yoruba-diacritic-restoration Automatic Yorùbá

ayomidedavid/yoruba-diacritic-restoration-model

Yoruba Diacritic Restoration — Hybrid BiLSTM + Transformer This project provides a scaffold to trai