Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

English-Giriama Parallel Sentence Dataset

Domain:

natural language processing

Record type:

dataset
Creator:
Lin
Host:
This dataset consists of sentence pairs in English and their corresponding translations in Giriama (Kigiryama), a Bantu language spoken primarily in coastal Kenya. It supports machine translation (MT) and other cross-lingual NLP tasks, especially in low-resource language research. Each row in the dataset contains: English Sentence: A sentence in standard English.

Visit

huggingface.co

Tasks

machine translation

Languages

Kigiryama

Tags

machine-translationlow-resourceafrican-languagesEnglishGiriamaagriculturebibleparallel-corpus

Licenses

cc-by-4.0

Similar

Sentence-aligned parallel corpus Amazigh-EnglishParallel English–Akan Navigation Dataset Parallel English–Akan Navigation DatasetTwi-English Parallel DatasetPristine Twi-English Parallel DatasetPristine Twi-English Parallel DatasetSinhala-English Parallel Word Dictionary Dataset

Sentence-aligned parallel corpus Amazigh-English

Parallel English–Akan Navigation Dataset Parallel English–Akan Navigation Dataset

This corpus was developed to support assistive navigation technology for visually impaired Akan-spea

Twi-English Parallel Dataset

This dataset contains a large-scale parallel corpus of Twi-English sentence pairs, featuring synthet

Pristine Twi-English Parallel Dataset

A large-scale Twi ↔ English parallel dataset derived from the Pristine Twi Dataset by the Ghana NLP

Pristine Twi-English Parallel Dataset

A large-scale Twi ↔ English parallel dataset derived from the Pristine Twi Dataset by the Ghana NLP

Sinhala-English Parallel Word Dictionary Dataset

Parallel datasets are vital for performing and evaluating any kind of multilingual task. However, in