Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

xri/KonongoNMT

Domain:

natural language processing

Record type:

dataset
Creator:
xri
Host:
Performance: BLEU 8.55, ChrF 50.372 KonongoNMT is a parallel dataset composed of 8,000 sentences in Swahili and Konongo. It is intended to be used for fine-tuning Neural Machine Translation models and Large Language Models for Konongo. Konongo is a low-resource Bantu language spoken in Tanzania.

Visit

huggingface.co

Tasks

machine translation

Languages

KonongoSwahili

Licenses

cc-by-sa-4.0

Similar

xri/NgoremeNMTxri/RwilaNMTxri/dagaare_synth_trainxri/songhay_synth_trainxri/dagaare_synth_source_targetxri/kanuri_synth_train

xri/NgoremeNMT

NgoremeNMT is a parallel dataset composed of 8,000 sentences in Swahili and Ngoreme. It is intended

xri/RwilaNMT

Performance: 15.26 BLEU, 54.53 ChrF RwilaNMT is a parallel dataset composed of 8,000 sentences in Sw

xri/dagaare_synth_train

The dagaare_dict_guided_train_9k.tsv presents the synthesized data for the dictionary-guided trainin

xri/songhay_synth_train

The songhay_dict_guided_train_9k.tsv presents the synthesized data for the dictionary-guided trainin

xri/dagaare_synth_source_target

The dagaareDictTrain.tsv was generated using the Machine Translation from One Book (MTOB) technique

xri/kanuri_synth_train

The kanuri_dict_guided_train_9k.tsv presents the synthesized data for the dictionary-guided training