This dataset is a compilation of Tamazight (zgh) language resources created by CIEMEN as part of the Awal project (
awaldigital.org), with funding from the Municipality of Barcelona and the Government of Catalonia. It includes 1,002 monolingual sentences from a Tamazight language learning material, and over 417,000 parallel sentence pairs spanning multiple language pairs: English–Tamazight, French–Tamazight, Catalan–Tamazight, Spanish–Tamazight, and Arabic–Tamazight. Parallel data comes from community contributions to the Awal platform, Tatoeba sentence pairs transliterated into Tifinagh script, Tamazight proverbs, web localization strings (Mozilla Common Voice, Awal platform), and segmented document translations. The dataset totals approximately 4.6 million words across all files.