Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Wolof Non-Standard to Standard Parallel Pairs

Domain:

natural language processing

Record type:

dataset
Creator:
soy
Host:
This dataset contains pairs of non-standard and standard Wolof text, designed for training models to normalize informal Wolof writing found on social media, messaging apps, and online platforms. The non-standard versions simulate real-world informal Wolof text with French code-switching, phonetic spellings, missing diacritics, and common typing variations.

Visit

huggingface.co

Tasks

text normalization

Languages

Wolof

Licenses

cc-by-sa-4.0

Similar

Afrikaans in a quantitative microtypology of Germanic standard and non-standard varietiesLearning to Explain Non-Standard English Words and PhrasesExtracting "non-standard" data from the Twitter APINormalization of Non-Standard Words in Croatian TextsThe Standard Wolof Morphology: A Descriptive Study of the Adjectives158 - ViDia2Std: A Parallel Corpus and Methods for Low-Resource Vietnamese Dialect-to-Standard Translation

Afrikaans in a quantitative microtypology of Germanic standard and non-standard varieties

Besides English, Afrikaans is considered “the [Germanic] language which deviates grammatically the f

Learning to Explain Non-Standard English Words and Phrases

We describe a data-driven approach for automatically explaining new, non-standard English expression

Extracting "non-standard" data from the Twitter API

The present paper examines methodology in the use of Twitter in the corpus-based
a

Normalization of Non-Standard Words in Croatian Texts

This paper presents text normalization which is an integral part of any text-to-speech synthesis sys

The Standard Wolof Morphology: A Descriptive Study of the Adjectives

This descriptive research paper analyzes the morphology of Wolof adjectives by distinguishing two ma

158 - ViDia2Std: A Parallel Corpus and Methods for Low-Resource Vietnamese Dialect-to-Standard Translation

Vietnamese exhibits extensive dialectal variation, posing challenges for NLP systems trained predomi