Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Building a Functional Machine Translation Corpus for Kpelle

Domain:

natural language processing

Record type:

paperdatasetmodel
Creator:
YamWeaDor
Host:avatar
In this paper, we introduce the first publicly available English-Kpelle dataset for machine translation, comprising over 2000 sentence pairs drawn from everyday communication, religious texts, and educational materials. By fine-tuning Meta's No Language Left Behind(NLLB) model on two versions of the dataset, we achieved BLEU scores of up to 30 in the Kpelle-to-English direction, demonstrating the benefits of data augmentation. Our findings align with NLLB-200 benchmarks on other African languages, underscoring Kpelle's potential for competitive performance despite its low-resource status. Beyond machine translation, this dataset enables broader NLP tasks, including speech recognition and language modelling. We conclude with a roadmap for future dataset expansion, emphasizing orthographic consistency, community-driven validation, and interdisciplinary collaboration to advance inclusive language technology development for Kpelle and other low-resourced Mande languages.

Visit

arxiv.org

Tasks

machine translation

Languages

MandinkaManinkakan, EasternNomaandeTobanga

Tags

Computation and Language

Similar

English-Twi Parallel Corpus for Machine TranslationA PARALLEL CORPUS FOR ADVANCING ENGLISH–SANTALI NEURAL MACHINE TRANSLATIONExtended Parallel Corpus for Amharic-English Machine TranslationOffline Corpus Augmentation for English-Amharic Machine TranslationAmharic-English Parallel Corpus for Neural Machine TranslationAmharic-English-Machine-Translation-Corpus

English-Twi Parallel Corpus for Machine Translation

We present a parallel machine translation training corpus for English and Akuapem Twi of 25,421 sentence pairs. We used a transformer-based translator to generate initial translations in Akuapem Twi, which were later verified and corrected where necessary by native

A PARALLEL CORPUS FOR ADVANCING ENGLISH–SANTALI NEURAL MACHINE TRANSLATION

Machine Translation (MT) poses a significant challenge in developing language corpora for low-resour

Extended Parallel Corpus for Amharic-English Machine Translation

This paper describes the acquisition, preprocessing, segmentation, and alignment of an Amharic-English parallel corpus. It will be helpful for machine translation of a low-resource language, Amharic. We freely released the corpus for research purposes. Furthermore,

Offline Corpus Augmentation for English-Amharic Machine Translation

The purpose of this study was to investigate the effect of corpus augmentation on the quality of English-Amharic Machine Translation (MT). In fact, trigram and four-gram Statistical Machine Translation (SMT) language models, as well as Neural Machine Translation (N

Amharic-English Parallel Corpus for Neural Machine Translation

Amharic is the working language of Ethiopia and, owing to its Semitic characteristics, the language

Amharic-English-Machine-Translation-Corpus

Amharic English Machine Translation Corpus prepared through website crawelling and custom preprocessing. This is a corpus made in effort to make amaharic english parallel data avilable for anyone who wants to deal with machine translation.