Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

ParaKab – Many Languages, One Kabyle: A Multilingual Parallel Corpus for a Low-Resource Language

Domaine:

natural language processing

Type de record:

dataset
Créateur:
MesTAGBouBad
Éditeur:
Zenodo
Hôte:avatar
Description of the Dataset This dataset consists of three parallel corpora involving the Kabyle language, covering the following language pairs: Kabyle-Arabic Kabyle-French Kabyle-English Together, these corpora comprise approximately one million (1M) aligned sentence pairs, forming a multilingual Kabyle-centric parallel dataset. The resource is intended for research in natural language processing (NLP), particularly for low-resource languages (LRLs), machine translation, and cross-linguistic studies. The data were compiled from multiple sources, primarily from publicly available resources distributed via the OPUS platform, as well as selected public websites. All texts underwent a rigorous pipeline including dataset selection, cleaning, normalization, alignment, verification, and deduplication, resulting in a ready-to-use dataset for research and machine learning applications. Academic Context This dataset was developed as part of a university research project on low-resource languages, focusing on Kabyle, a northern Tamazight (Berber) language. The project aims to contribute to the digital inclusion of low-resource languages and to support their integration into modern NLP systems. Licensing and Data Rights Each sentence pair in the dataset is linked to its original data source and, where applicable, its corresponding license or rights holder, ensuring transparency and compliance with source-specific terms of use. For the publication of this corpus, a general license is provided for the dataset as a whole. Users are responsible for respecting the original licenses and usage conditions associated with each source when reusing the data.

Visit

doi.orgzenodo.org

Tasks

machine translation

Languages

AmazighBerberGhomaraSenhaja BerberTamazight, Central AtlasTarifit

Tags

Low-Resource LanguagesMultilingual Parallel CorpusKabyleTamazightMachine TranslationNatural Language Processing

Licenses

Creative Commons Attribution 4.0 Internationalhttps://creativecommons.org/licenses/by/4.0/legalcodeProvenance and Source LicensesAll rights are retained by the creators, except when individual pairs have original source licenses or rights holders, as indicated in the metadata.http://rightsstatements.org/vocab/InC/1.0/