Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Awal Tamazight Dataset

Domain:

natural language processing

Record type:

dataset
Host:
This dataset is a compilation of Tamazight (zgh) language resources created by CIEMEN as part of the Awal project (awaldigital.org), with funding from the Municipality of Barcelona and the Government of Catalonia. It includes 1,002 monolingual sentences from a Tamazight language learning material, and over 417,000 parallel sentence pairs spanning multiple language pairs: English–Tamazight, French–Tamazight, Catalan–Tamazight, Spanish–Tamazight, and Arabic–Tamazight. Parallel data comes from community contributions to the Awal platform, Tatoeba sentence pairs transliterated into Tifinagh script, Tamazight proverbs, web localization strings (Mozilla Common Voice, Awal platform), and segmented document translations. The dataset totals approximately 4.6 million words across all files.

Visit

mozilladatacollective.com

Tasks

machine translation

Languages

AmazighBerberGhomaraSenhaja BerberTamazight, Central AtlasTamazight, Standard MoroccanTarifit

Tags

mdcmozilla data collectiveLMTSVJSONTXT

Licenses

Creative Commons Attribution 4.0 International (CC-BY-4.0)

Similar

Awal -- Community-Powered Language Technology for TamazightTamazight Numbers DatasetTamazight Open Speech DatasetTODa: Tamazight Open DatasetTamazight Open Speech DatasetAbdoussamadAgAlhousseini/awal

Awal -- Community-Powered Language Technology for Tamazight

This paper presents Awal, a community-powered initiative for developing language technology resource

Tamazight Numbers Dataset

This dataset contains numbers from 1 to 1,000,000 translated into: English. French. Spanish. Tamazi

Tamazight Open Speech Dataset

This dataset provides a parsed, formatted, and ready-to-use Amazigh Voice Dataset. It contains voice

TODa: Tamazight Open Dataset

Welcome to the Tamazight Open Dataset (TODa), a groundbreaking open-source project dedicated to pres

Tamazight Open Speech Dataset

This dataset provides a parsed, formatted, and ready-to-use Amazigh Voice Dataset. It contains voice

AbdoussamadAgAlhousseini/awal

Awal (ⴰⵡⴰⵍ) — plateforme mondiale de référence de la langue tamasheq. MVP dictionnaire v0.1, triling