Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Machine Translation for Kusaal: A Parallel Corpus and a First Open-Source System

Domain:

natural language processing

Record type:

datasetmodel
Creator:
Alh
Publisher:
Zenodo
Host:avatar
Kusaal is a Mabia (Gur) language spoken by roughly 420,000 people in Ghana's Upper East Region and across borders in Burkina Faso and Togo. Despite having a reference grammar, an official orthography, and a complete Bible translation, it has no machine translation system of any kind. This paper presents a 34,568-pair Kusaal–English parallel corpus assembled from six sources (YouVersion Bible, English–Kusaal Index, GhanaNLP, Lexique Pro, Wikipedia, and back-translation), and a bidirectional translation model fine-tuned from NLLB-200-distilled-600M. A new language token (kus_Latn) is initialised from the Dagbani (dag_Latn) embedding, the closest Mabia relative already present in NLLB, rather than from random noise. The model achieves 27.57 BLEU translating Kusaal into English and 13.72 BLEU in the reverse direction. The paper also documents two silent data corruptions encountered during corpus construction: an HTML leak in scraped Bible text, and a CSV quoting failure that silently dropped rows. Both are reported as a cautionary note for low-resource corpus work where every sentence pair is expensive. Corpus and model are released publicly under CC BY 4.0.

Visit

doi.org

Tasks

machine translation

Languages

DagbaniKusaal

Tags

kusaalNLPAfrican LanguagesMabia LanguagesGurNLLBGhanaBawkuKusaal, machine translation, low-resource NLP, African languages, Mabia languages, Gur languages, parallel corpus, NLLB, Ghana, NLP

Licenses

Creative Commons Attribution 4.0 Internationalhttps://creativecommons.org/licenses/by/4.0/legalcode