This dataset contains parallel sentences in English and Kabyle, cleaned and filtered using the GlotLid model. The dataset is derived from the OPUS-NLLB corpus and has been processed to ensure high-quality sentence pairs.
nllb_en_kab.parquet: A Parquet file containing the cleaned English-Kabyle sentence pairs.
Total Sentence Pairs: 2,484,297
English Sentences: 2,484,297