Aligned parallel corpora for the following language pair: English-SiSwati. The data is given as four
Monolingual corpus for SiSwati. The data is given as a single UTF-8 text file, with each segment on
NCHLT corpus of morphologically annotated tokens in Siswati converted to the tags used during phases