Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Bilingual English-Siswati Corpus

Domain:

natural language processing

Record type:

dataset
Creator:
McKellar, Cindy
Publisher:
North-West University - Centre for Text Technology (CTexT)
Host:avatar
Aligned parallel corpora for the following language pair: English-SiSwati. The data is given as four separate UTF-8 text files, with each segment on a newline. Dataset contains existing data sourced for the DSAC funded Autshumato project as well as new data sourced for the SADiLaR: Parallel corpora for English into SiSwati project. The dataset contains the following types of bilingual data: Translations from English to Siswati and crawled parallel data for English-Siswati. The dataset comprises a total of 114,839 segments with 2,002,293 English words and 1, 423,414 SiSwati words.

Visit

hdl.handle.net

Tasks

machine translation

Languages

Swati

Tags

Siswati, aligned data, multilingual, translations, crawled, machine translation training data

Licenses

Creative Commons Attribution 4.0 International

Similar

Bilingual English-isiXhosa corpusAmharic-English bilingual corpusWEB MINING FOR AN AMHARIC - ENGLISH BILINGUAL CORPUSSiswati Ner CorpusMonolingual Siswati CorpusSiswati NER Corpus

Bilingual English-isiXhosa corpus

Aligned parallel corpora for the following language pair: English-isiXhosa. The data is given as tw

Amharic-English bilingual corpus

The Amharic-English bilingual corpus contains parallel text from legal and news domains in Amharic s

WEB MINING FOR AN AMHARIC - ENGLISH BILINGUAL CORPUS

Siswati Ner Corpus

Named entity annotated data from the NCHLT Text Resource Development: Phase II Project, annotated with PERSON, LOCATION, ORGANISATION and MISCELLANEOUS tags.

Monolingual Siswati Corpus

Monolingual corpus for SiSwati. The data is given as a single UTF-8 text file, with each segment on

Siswati NER Corpus

Named entity annotated data from the NCHLT Text Resource Development: Phase II Project, annotated wi