Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

English-Kabyle Parallel Corpus (translatewiki)

Domain:

natural language processing

Record type:

dataset
Creator:
bof
Host:
A clean, deduplicated parallel corpus of English → Kabyle (Taqbaylit) translations extracted from the translatewiki.net bulk dump (2026-01-01). Attribute Value Language pair English (en) → Kabyle (kab) Total unique pairs 8,871 Source translatewiki.net License CC BY 3.0 Domain Software localization, UI strings, documentation { "translation": {

Visit

huggingface.co

Tasks

machine translation

Languages

Amazigh

Tags

kabyletaqbaylitlow-resourcewikimediaberbertamazighttamaziɣt

Licenses

cc-by-3.0

Similar

Imsidag-community/english-kabyle-parallelabdelhaqueidali/Kabyle-Latin-to-Tifinagh-Parallel-CorpusAkan–English Maternal Health Parallel Corpus Akan–English Maternal Health Parallel CorpusSomali-English Parallel CorpusNupe-English parallel corpusAmharic-English Parallel Corpus

Imsidag-community/english-kabyle-parallel

130 883 aligned sentence pairs extracted from the open Tatoeba database. Download – raw Tatoeba dum

abdelhaqueidali/Kabyle-Latin-to-Tifinagh-Parallel-Corpus

This dataset provides a parallel corpus of the Kabyle language (Taqbaylit), pairing native Latin-bas

Akan–English Maternal Health Parallel Corpus Akan–English Maternal Health Parallel Corpus

This dataset contains a curated bilingual parallel corpus developed to support domain-specific neura

Somali-English Parallel Corpus

This dataset contains high-quality parallel sentence pairs, multi-sentence alignments, and paragraph

Nupe-English parallel corpus

This is the first ever Nupe - English Parallel Corpus and Nupe Monolingual Corpora curated from diverse sources including poems,idioms, proverbs, religpoius text etc. The aim of this data collection is to make available a cultural-aware Nupe-english corpus for NLP

Amharic-English Parallel Corpus

This corpus consists of 145,820 Amharic-English parallel sentences (segments) from various sources. This corpus is larger in size than previously compiled corpora. It is released for research purposes and can be used to train or support Amharic-English machine tran