Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

ErwinSchillack/Comon-Voice-Extraction-Data-Collection-South-African-Languages

Domain:

natural language processing

Record type:

software
Creator:
Erw
Host:
Collection and extraction of data from South African languages # Comon-Voice-Extraction-Data-Collection-South-African-Languages Collection and extraction of data from South African languages cv-sentence-extractor is a Rust library typically used in Artificial Intelligence, Natural Language Processing applications Common Voice is Mozilla's initiative to help teach machines how real people speak. For this we need to collect sentences that people can read out aloud on the website. Individual sentences can be submitted through the Sentence Collector. This only can scale so far, so we also use automated tools to extract sentences from other sources. Words and sentences were extracted from wikipedai related to the 11 official South African Languages. CText NCHLT Web Service is then used to indetify the given language. # Setup Clone this repo: ```bash git clone github.com ``` Next, download the WikiExtractor: ```bash git clone github.com ``` # Extraction Type the following in the terminal and change the XX to the language that it corresponds to. Ex: en for English, af for Afrikaans and ve for Tshivenḓa and so on. ```bash wget dumps.wikimedia.org bzip2 -d XXwiki-latest-pages-articles-multistream.xml.bz2 ``` Use WikiExtractor to extract the dump. In the parameters, we specify to use JSON as the output format instead of the default XML. ```bash cd wikiextractor git checkout e4abb4cbd019b0257824ee47c23dd163919b731b python WikiExtractor.py --json ../XXwiki-latest-pages-articles-multistream.xml ``` # Rule Files After extracting the dump, we need to setup a rules.otml file for the specific language that we are working with. | Name | Description | Values | Default | |--------|-----------------------|---------|---------| | abbreviation_patterns | Regex defining abbreviations | Rust Regex Array | all abbreviations allowed | allowed_symbols_regex | Regex of allowed symbo …

Visit

github.com

Languages

AfrikaansVenda

Similar

TWB Voice Playbook for voice data collection for low-resource languagesCommon Voice (African Languages)Synthetic Voice Data for Automatic Speech Recognition in African LanguagesEmbedding Evaluation Data for South African LanguagesSouth African National Diatom Collection (SANDC)On Extraction, argument binding and voice morphology in Malagasy

TWB Voice Playbook for voice data collection for low-resource languages

This playbook will help you to plan and manage projects to collect voice data for low-resource languages. It is aimed at both new and experienced teams and covers the full process, from setting up the project to publishing your dataset. We draw on CLEAR Global’s

Common Voice (African Languages)

Speech recognition, voice assistants Notes / challenges: Requires cleaning for production use

Synthetic Voice Data for Automatic Speech Recognition in African Languages

Speech technology remains out of reach for most of the over 2300 languages in Africa. We present the first systematic assessment of large-scale synthetic voice corpora for African ASR.

We apply a three-step process: LLM-driven text creation, TTS voice sy

Embedding Evaluation Data for South African Languages

WordSim and Simlex Data for South African Languages

South African National Diatom Collection (SANDC)

The South African National Diatom Collection (SANDC) contains many thousands of records pertaining t

On Extraction, argument binding and voice morphology in Malagasy