Logo Lanfrica

kimrichies/English-Luganda-Parallel-corpus

Domain:

natural language processing

Record type:

dataset
Creator:
kim
Host:
This is a bilingual corpus of English and Luganda for use in Neural Machine Translation tasks. I give credit to Zenodo and Sunbird AI for the public datasets that gave a foundation to this work # English-Luganda-Parallel-corpus This is a bilingual corpus of English and Luganda for use in Neural Machine Translation tasks. I give credit to Zenodo and Sunbird AI for the public datasets that gave a foundation to this work. I have also included the King James version of the bible text that is well alighned between English and Luganda, this was extracted from [Bible World Project] (wordproject.org) We used this dataset to build custom Neural Machine Translation models for Luganda and English. After hyperparameter tuning, we achieved a BLEU score of 21.28 for English-to-Luganda and 17.47 for Luganda-to-English. Reference papers; Kimera, R., Rim, D.N. and Choi, H., 2022. Building a Parallel Corpus and Training Translation Models Between Luganda and English. KIISE, 49(11), pp.1009-1016. Rim, D. N., Kimera, R., & Choi, H. (2023). Mini-Batching with Similar-Length Sentences to Quickly Train NMT Models. KIISE, 50(7), 614-620. Acknowledgement; [MILAb] (milab.handong.edu)