Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Twi-English Parallel Dataset

Domain:

natural language processing

Record type:

dataset
Creator:
mic
Host:
This dataset contains a large-scale parallel corpus of Twi-English sentence pairs, featuring synthetically generated Twi sentences with corresponding English paraphrases. The dataset is designed to support machine translation, cross-lingual understanding, and other NLP tasks involving the Twi language (a dialect of Akan spoken in Ghana). Total Size: 47,924,398 parallel sentence pairs

Visit

huggingface.co

Tasks

machine translation

Languages

AkanBwamu, CwiDinka, SoutheasternTwi

Tags

twiakanafrican-languageslow-resourcesynthetic-dataparallel-corpus

Licenses

cc-by-4.0