Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

AnthonyEzra/english-tumbuka-parallel-corpus

Domain:

natural language processing

Record type:

dataset
Creator:
Ant
Host:
This repository contains parallel sentence data for English ↔ Tumbuka translation tasks. It includes both training and evaluation splits in CSV format. Purpose: Training Size: 232,536 sentence pairs Content: Mixed-quality bilingual pairs. Source: Partially adapted from michsethowusu/english-tumbuka_sentence-pairs_mt560

Visit

huggingface.co

Tasks

machine translation

Languages

Tumbuka

Similar

AnthonyEzra/english-tumbuka-bibleAkan–English Maternal Health Parallel Corpus Akan–English Maternal Health Parallel CorpusSomali-English Parallel CorpusNupe-English parallel corpusAmharic-English Parallel Corpusluganda-english-parallel-corpus

AnthonyEzra/english-tumbuka-bible

Akan–English Maternal Health Parallel Corpus Akan–English Maternal Health Parallel Corpus

This dataset contains a curated bilingual parallel corpus developed to support domain-specific neura

Somali-English Parallel Corpus

This dataset contains high-quality parallel sentence pairs, multi-sentence alignments, and paragraph

Nupe-English parallel corpus

This is the first ever Nupe - English Parallel Corpus and Nupe Monolingual Corpora curated from diverse sources including poems,idioms, proverbs, religpoius text etc. The aim of this data collection is to make available a cultural-aware Nupe-english corpus for NLP

Amharic-English Parallel Corpus

This corpus consists of 145,820 Amharic-English parallel sentences (segments) from various sources. This corpus is larger in size than previously compiled corpora. It is released for research purposes and can be used to train or support Amharic-English machine tran

luganda-english-parallel-corpus