Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Text Normalization for Low-Resource Languages of Africa

Domain:

natural language processing

Record type:

paper
Training data for machine learning models can come from many different sources, which can be of dubious quality. For resource-rich languages like English, there is a lot of data available, so we can afford to throw out the dubious data. For low-resource languages where there is much less data available, we can't necessarily afford to throw out the dubious data, in case we end up with a training set which is too small to train a model. In this study, we examine the effects of text normalization and data set quality for a set of low-resource languages of Africa -- Afrikaans, Amharic, Hausa, Igbo, Malagasy, Somali, Swahili, and Zulu. We describe our text normalizer which we built in the Pynini framework, a Python library for finite state transducers, and our experiments in training language models for African languages using the Natural Language Toolkit (NLTK), an open-source Python library for NLP.

Visit

arxiv.org

Connected records

project

Tasks

text normalization

Languages

AfrikaansAmharicHausaIgboMalagasy, MerinaSomaliSwahiliZulu

Similar

Text Normalization for Low Resource Languagesgoogleinterns/text-norm-for-low-resource-languagesMassively Multilingual Text Translation For Low-Resource LanguagesAdversarial Text-to-Speech for low-resource languagesXF2T: Cross-lingual Fact-to-Text Generation for Low-Resource LanguagesText Image Generation for Low-Resource Languages with Dual Translation Learning

Text Normalization for Low Resource Languages

This repository contains code related to the Google open source internship project Text Normalization for Low Resource Languages.

googleinterns/text-norm-for-low-resource-languages

# Text Normalization for Low Resource Languages This repository contains code related to the Google

Massively Multilingual Text Translation For Low-Resource Languages

Translation into severely low-resource languages has both the cultural goal of saving and reviving t

Adversarial Text-to-Speech for low-resource languages

Improving the adversarial TTS models for low-resource languages by utilizing the high-frequency similarities between the different languages.

XF2T: Cross-lingual Fact-to-Text Generation for Low-Resource Languages

Multiple business scenarios require an automated generation of descriptive human-readable text from

Text Image Generation for Low-Resource Languages with Dual Translation Learning

Scene text recognition in low-resource languages frequently faces challenges due to the limited avai