Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Autshumato English-Siswati Parallel Corpora

Domain:

natural language processing

Record type:

dataset
Creator:
McKellar, Cindy
Publisher:
North-West University - Centre for Text Technology (CTexT)
Host:avatar
Aligned parallel corpora for the following language pair: English-SiSwati. The data is given as four separate UTF-8 text files, with each segment on a newline. Dataset contains existing data sourced for the DSAC funded Autshumato project as well as new data sourced for the SADiLaR: Parallel corpora for English into SiSwati project. The dataset contains the following types of bilingual data: Translations from English to Siswati and crawled parallel data for English-Siswati. The dataset comprises a total of 114,839 segments with 2,002,293 English words and 1, 423,414 SiSwati words. (A new version issued since the title was changed)

Visit

hdl.handle.net

Tasks

machine translation

Languages

Swati

Tags

Siswatimachine translation training datacrawledtranslationsmultilingualaligned data

Licenses

Creative Commons Attribution 4.0 International

Similar

Autshumato English-Sepedi Parallel CorporaAutshumato English-isiZulu Parallel CorporaAutshumato English-Setswana Parallel CorporaAutshumato English-Tshivenḓa Parallel CorporaAutshumato English-Afrikaans Parallel CorporaAutshumato English-Sesotho Parallel Corpora

Autshumato English-Sepedi Parallel Corpora

Aligned parallel corpora for the language pair English-Sepedi. The data is given as two separate UTF

Autshumato English-isiZulu Parallel Corpora

Aligned parallel corpora for the language pair English-isiZulu. The data is given as two separate UT

Autshumato English-Setswana Parallel Corpora

Aligned English-Setswana parallel corpus. This set contains data that was translated by professional

Autshumato English-Tshivenḓa Parallel Corpora

Aligned parallel corpora for the following language pair: English-Tshivenḓa. Data was crawled from v

Autshumato English-Afrikaans Parallel Corpora

Aligned parallel corpora for the language pair English-Afrikaans. The data is given as two separate

Autshumato English-Sesotho Parallel Corpora

Aligned parallel corpora for the language pair English-Sesotho. The data is given as two separate UT