Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Autshumato English-Siswati Parallel Corpora

Domaine:

natural language processing

Type de record:

dataset
Créateur:
McKellar, Cindy
Éditeur:
North-West University - Centre for Text Technology (CTexT)
Hôte:avatar
Aligned parallel corpora for the following language pair: English-SiSwati. The data is given as four separate UTF-8 text files, with each segment on a newline. Dataset contains existing data sourced for the DSAC funded Autshumato project as well as new data sourced for the SADiLaR: Parallel corpora for English into SiSwati project. The dataset contains the following types of bilingual data: Translations from English to Siswati and crawled parallel data for English-Siswati. The dataset comprises a total of 114,839 segments with 2,002,293 English words and 1, 423,414 SiSwati words. (A new version issued since the title was changed)

Visit

hdl.handle.net

Tasks

machine translation

Languages

Swati

Tags

Siswatimachine translation training datacrawledtranslationsmultilingualaligned data

Licenses

Creative Commons Attribution 4.0 International

Similaires

Autshumato English-Xitsonga Parallel CorporaAutshumato English-Setswana Parallel CorporaAutshumato English-Afrikaans Parallel CorporaAutshumato English-Sepedi Parallel CorporaAutshumato English-isiZulu Parallel CorporaAutshumato English-Sesotho Parallel Corpora

Autshumato English-Xitsonga Parallel Corpora

Aligned English-Xitsonga parallel corpus. The data is given as two seperate UTF-8 text files; with e

Autshumato English-Setswana Parallel Corpora

Aligned English-Setswana parallel corpus. This set contains data that was translated by professional

Autshumato English-Afrikaans Parallel Corpora

Aligned parallel corpora for the language pair English-Afrikaans. The data is given as two separate

Autshumato English-Sepedi Parallel Corpora

Aligned parallel corpora for the language pair English-Sepedi. The data is given as two separate UTF

Autshumato English-isiZulu Parallel Corpora

Aligned parallel corpora for the language pair English-isiZulu. The data is given as two separate UT

Autshumato English-Sesotho Parallel Corpora

Aligned parallel corpora for the language pair English-Sesotho. The data is given as two separate UT