Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

TODa: Tamazight Open Dataset

Domaine:

natural language processing

Type de record:

dataset
Hôte:
Welcome to the Tamazight Open Dataset (TODa), a groundbreaking open-source project dedicated to preserving and advancing the Tamazight language. With its extensive collection of linguistic data, TODa stands as a pioneering collaborative project for Tamazight <=> Englis translation, specifically designed for Natural Language Processing applications. TODa's unique approach combines both semantic and syntactic categorization methods, offering a rich representation of words in their various contexts and forms. The dataset encompasses a comprehensive collection of linguistic elements, including detailed verb conjugations across different tenses, noun variations, and an extensive compilation of translated expressions that capture the language's nuances. What sets TODa apart is its inclusive approach to Tamazight's writing systems. The dataset thoughtfully incorporates Latin alphabets, acknowledging and preserving the diverse writing traditions practiced across Amazigh communities. This dual-script approach ensures broader accessibility and cultural authenticity. Our vision is to establish TODa as the cornerstone resource for Tamazight Natural Language Processing. Through this meticulously curated dataset, we strive to empower developers and researchers to create innovative NLP solutions that authentically serve the Amazigh-speaking community. We take pride in our current progress, yet acknowledge that language documentation is an evolving journey. We actively encourage participation from the Amazigh technology community to contribute their expertise in expanding and refining the dataset. Through collaborative effort, we can create a robust foundation for technological innovations that honor and advance Amazigh linguistic heritage.

Visit

mozilladatacollective.com

Tasks

machine translation

Languages

AmazighBerberGhomaraSenhaja BerberTamazight, Central AtlasTamazight, Standard MoroccanTarifitTedaga

Tags

mdcmozilla data collectiveNLPCSV

Licenses

Creative Commons Attribution 4.0 International (CC-BY-4.0)

Similaires

Tamazight Open Speech DatasetTamazight Open Speech DatasetOpen-Darija-Tamazight/tamazight-english-translateTamazight Numbers DatasetAwal Tamazight DatasetOpen-Darija-Tamazight/chat-darija

Tamazight Open Speech Dataset

This dataset provides a parsed, formatted, and ready-to-use Amazigh Voice Dataset. It contains voice

Tamazight Open Speech Dataset

This dataset provides a parsed, formatted, and ready-to-use Amazigh Voice Dataset. It contains voice

Open-Darija-Tamazight/tamazight-english-translate

Tamazight Numbers Dataset

This dataset contains numbers from 1 to 1,000,000 translated into: English. French. Spanish. Tamazi

Awal Tamazight Dataset

This dataset is a compilation of Tamazight (zgh) language resources created by CIEMEN as part of the

Open-Darija-Tamazight/chat-darija