Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Attentive Sequence-to-Sequence Learning for Diacritic Restoration of Yorùbá Language Text

Domain:

natural language processing

Record type:

paper
Yorùbá is a widely spoken West African language with a writing system rich in tonal and orthographic diacritics. With very few exceptions, diacritics are omitted from electronic texts, due to limited device and application support. Diacritics provide morphological information, are crucial for lexical disambiguation, pronunciation and are vital for any Yorùbá text-to-speech (TTS), automatic speech recognition (ASR) and natural language processing (NLP) tasks. Reframing Automatic Diacritic Restoration (ADR) as a machine translation task, we experiment with two different attentive Sequence-to-Sequence neural models to process undiacritized text. On our evaluation dataset, this approach produces diacritization error rates of less than 5%. We have released pre-trained models, datasets and source-code as an open-source project to advance efforts on Yorùbá language technology.

Visit

arxiv.orgwww.isca-speech.org

Connected records

model

Tasks

diacritic restorationtext normalization

Languages

Yoruba

Similar

Automatic Diacritic Restoration of Yorùbá language TextImproving Yorùbá Diacritic RestorationMultilingual sequence-to-sequence speech recognition: architecture, transfer learning, and language modelingDiacritic Restoration for Yoruba Text with under dot and Diacritic Mark Based on LSTMDiacritic Restoration in Yoruba LanguageLanguage model integration based on memory control for sequence to sequence speech recognition

Automatic Diacritic Restoration of Yorùbá language Text

Automatic Diacritic Restoration of Yorùbá language Text

Improving Yorùbá Diacritic Restoration

Yorùbá is a widely spoken West African language with a writing system rich in orthographic and tonal diacritics. They provide morphological information, are crucial for lexical disambiguation, pronunciation and are vital for any computational Speech or Natural Lang

Multilingual sequence-to-sequence speech recognition: architecture, transfer learning, and language modeling

Sequence-to-sequence (seq2seq) approach for low-resource ASR is a relatively new direction in speech

Diacritic Restoration for Yoruba Text with under dot and Diacritic Mark Based on LSTM

Abstract Yoruba is a tonal language spoken primarily in Nigeria, some West African countries,

Diacritic Restoration in Yoruba Language

A project on diacritic restoration of languages

Language model integration based on memory control for sequence to sequence speech recognition

In this paper, we explore several new schemes to train a seq2seq model to integrate a pre-trained LM