Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Niger-Volta-LTI/yoruba-adr

Domain:

natural language processing

Record type:

softwaremodel
Creator:
Nig
Host:
Automatic Diacritic Restoration of Yorùbá language Text # Automatic Diacritic Restoration of Yorùbá Text ### Motivations Nigeria’s dying languages! ### Applications * Generating very large, high quality Yorùbá text corpora * [physical books] → OCR → [undiacritized text] → ADR → [clean diacritized text] Physical books written in Yorùbá (novels, manuals, school books, dictionaries) are digitized via Optical Character Recognition (OCR), which may not fully respect tonal or orthographic diacritics. Next, the undiacritized text is processed to restore the correct diacritics. * Correcting digital texts scraped online on Twitter, Naija forums, articles, etc * Suggesting corrections during manual text entry (spell/diacritic checker) * Preprocessing text for training Yorùbá * language models * word embeddings * text-language identification (so Twitter can stop claiming Yorùbá text is Vietnamese haba!) * part-of-speech taggers * named-entity recognition * text-to-speech (TTS) models (speech synthesis) * speech-to-text (STT) models (speech recogntion) ### Pretrained ADR Models * New Pretrained Models (April 2020) → Evaluation Results * Older soft-attention models (March/June 2019) ### Datasets Niger Volta LTI: Yorùbá lan… ## Train a Yorùbá ADR model ### Dependencies * Python3 (tested on 3.5, 3.6, 3.7) * Install all dependencies: `pip3 install -r requirements.txt` We train models on an Amazon EC2 `p2.xlarge` instance running `Deep Learning AMI (Ubuntu) Version 5.0 (ami-c27af5ba)`. These machine-images (AMI) have Python3 and PyTorch pre-installed as well as CUDA for training on the GPU. We use the OpenNMT-py framework for training and restoration. * To install PyTorch 0.4 manually, follow instructions for your {OS, package manager, python, CUDA} versions * `git clone github.com` * `git clone github.com` * Install dependencies: `pip3 install -r requirements.txt` * Note that NLTK will need some extra hand-holding if you've in …

Visit

github.com

Tasks

diacritic restorationtext normalization

Languages

Yoruba

Tags

adrafrican-languagesattentiondiacriticsneural-machine-translationorthographic-diacriticspython3seq2seqtext-processingyoruba

Licenses

MIT

Similar

Niger-Volta-LTI/yoruba-textNiger-Volta-LTI/yoruba-voiceNiger-Volta-LTI/yoruba-audioNiger Volta LTI: Yoruba AudioNiger-Volta-LTI/yoruba-voice-speech-recorderNiger-Volta-LTI/Niger-Volta-LTI.github.io

Niger-Volta-LTI/yoruba-text

Yorùbá language training text for NLP, ASR and TTS tasks # Yorùbá text This repository contains fu

Niger-Volta-LTI/yoruba-voice

Repo & Project for the Imminent Research Grant code & tasks # YorùbáVoice Landing page for data, c

Niger-Volta-LTI/yoruba-audio

Yorùbá language audio for ASR, TTS and other speech tasks # Yorùbá Audio This repo aggregates audi

Niger Volta LTI: Yoruba Audio

This repo aggregates audio/speech corpora for Yorùbá tasks. The corpora may contain aligned text or be purely unlabeled. The objective is to have a bird's eye view of available Yorùbá audio, and it's metadata and entropy, to inform additional data collection tasks

Niger-Volta-LTI/yoruba-voice-speech-recorder

App for recording speech utterances dictated from text prompts. Speaker name, audio-recording path &

Niger-Volta-LTI/Niger-Volta-LTI.github.io

blog ## Blog @ Niger-Volta-LTI.github.io Ẹ káàbọ̀! Kwabɔ! Ób’ókhían! Wá do! Welcome! Bienvenue! to