Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Corpus-Based Approaches to Igbo Diacritic Restoration

Domain:

natural language processing

Record type:

paper
Creator:
Eze
Host:avatar
With natural language processing (NLP), researchers aim to enable computers to identify and understand patterns in human languages. This is often difficult because a language embeds many dynamic and varied properties in its syntax, pragmatics and phonology, which need to be captured and processed. The capacity of computers to process natural languages is increasing because NLP researchers are pushing its boundaries. But these research works focus more on well-resourced languages such as English, Japanese, German, French, Russian, Mandarin Chinese, etc. Over 95% of the world's 7000 languages are low-resourced for NLP, i.e. they have little or no data, tools, and techniques for NLP work. In this thesis, we present an overview of diacritic ambiguity and a review of previous diacritic disambiguation approaches on other languages. Focusing on the Igbo language, we report the steps taken to develop a flexible framework for generating datasets for diacritic restoration. Three main approaches, the standard n-gram model, the classification models and the embedding models were proposed. The standard n-gram models use a sequence of previous words to the target stripped word as key predictors of the correct variants. For the classification models, a window of words on both sides of the target stripped word was used. The embedding models compare the similarity scores of the combined context word embeddings and the embeddings of each of the candidate variant vectors. 270 page. Ph.D. Thesis. The University of Sheffield

Visit

arxiv.org

Tasks

diacritic restorationtext normalization

Languages

Igbo

Tags

thesisComputers and SocietyInformation Retrievaldiacriticsuniversity of sheffield

Similar

Lexical Disambiguation of Igbo using Diacritic RestorationCrinmatic/Diacritic-RestorationImproving Yorùbá Diacritic RestorationNatashadonoh11/yoruba-diacritic-restorationmajtom91/yoruba-diacritic-restorationDiacritic Restoration for Yoruba Text with under dot and Diacritic Mark Based on LSTM

Lexical Disambiguation of Igbo using Diacritic Restoration

Properly written texts in Igbo, a low-resource African language, are rich in both orthographic and tonal diacritics. Diacritics are essential in capturing the distinctions in pronunciation and meaning of words, as well as in lexical disambiguation. Unfortunately, m

Crinmatic/Diacritic-Restoration

Using AI to restore Diacritics on Yoruba language (which is a low resource language) ## **Diacritic

Improving Yorùbá Diacritic Restoration

Yorùbá is a widely spoken West African language with a writing system rich in orthographic and tonal diacritics. They provide morphological information, are crucial for lexical disambiguation, pronunciation and are vital for any computational Speech or Natural Lang

Natashadonoh11/yoruba-diacritic-restoration

Yoruba diacritic restoration- Adaption Labs submission # Yorùbá Diacritic Restoration Adapting a l

majtom91/yoruba-diacritic-restoration

Automatic Yorùbá diacritic restoration application # Yoruba-diacritic-restoration Automatic Yorùbá

Diacritic Restoration for Yoruba Text with under dot and Diacritic Mark Based on LSTM

Abstract Yoruba is a tonal language spoken primarily in Nigeria, some West African countries,