Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Siswati NER Corpus

Domain:

natural language processing

Record type:

dataset
Creator:
nwu
Host:
Named entity annotated data from the NCHLT Text Resource Development: Phase II Project, annotated with PERSON, LOCATION, ORGANISATION and MISCELLANEOUS tags.

Visit

huggingface.co

Tasks

information extractionnamed entity recognition

Languages

Swati

Licenses

other

Similar

Siswati Ner CorpusMonolingual Siswati CorpusNCHLT Speech Corpus -- siSwatiLwazi Siswati TTS corpusAutshumato Monolingual Siswati CorpusBilingual English-Siswati Corpus

Siswati Ner Corpus

Named entity annotated data from the NCHLT Text Resource Development: Phase II Project, annotated with PERSON, LOCATION, ORGANISATION and MISCELLANEOUS tags.

Monolingual Siswati Corpus

Monolingual corpus for SiSwati. The data is given as a single UTF-8 text file, with each segment on

NCHLT Speech Corpus -- siSwati

This is the siSwati language part of the NCHLT Speech Corpus of the South African languages. Languag

Lwazi Siswati TTS corpus

Orthographic and phonemically aligned transcriptions

Autshumato Monolingual Siswati Corpus

Monolingual corpus for SiSwati. The data is given as a single UTF-8 text file, with each segment on

Bilingual English-Siswati Corpus

Aligned parallel corpora for the following language pair: English-SiSwati. The data is given as four