Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Tamazight-Arabic Speech Translation Dataset

Domain:

natural language processing

Record type:

dataset
Creator:
Tam
Host:
This is the Tamazight-NLP organization-hosted version of the Tamazight-Arabic Speech Recognition Dataset. This dataset contains ~15.5 hours of Tamazight (Tachelhit dialect) speech paired with Arabic transcriptions, designed for automatic speech recognition (ASR) and speech-to-text translation tasks. Total Examples: 20,344 audio segments Training Set: 18,309 examples (~8.9GB)

Visit

huggingface.co

Tasks

automatic speech recognitionmachine translationspeech processingspeech translation

Languages

AmazighBerberGhomaraSenhaja BerberTachelhitTamazight, Central AtlasTamazight, Standard MoroccanTarifit

Tags

speechtamazightarabicspeech-to-textlow-resourcenorth-africalarge datasets from Lanfrica Insights

Similar

Tamazight-Arabic Speech Recognition DatasetTamazight-Arabic Speech Recognition Datasetaitdihimnassim/Tamazight-Arabic-TranslationA Novel Dataset for Arabic Speech Recognition Recorded by Tamazight SpeakersTamazight Open Speech DatasetTamazight Open Speech Dataset

Tamazight-Arabic Speech Recognition Dataset

This dataset contains speech segments in Tamazight (specifically focusing on the Tachelhit dialect)

Tamazight-Arabic Speech Recognition Dataset

This is the EMINES organization-hosted version of the Tamazight-Arabic Speech Recognition Dataset, s

aitdihimnassim/Tamazight-Arabic-Translation

A Novel Dataset for Arabic Speech Recognition Recorded by Tamazight Speakers

Automatic Speech Recognition (ASR) is an area of research that's constantly evolving, thanks to impo

Tamazight Open Speech Dataset

This dataset provides a parsed, formatted, and ready-to-use Amazigh Voice Dataset. It contains voice

Tamazight Open Speech Dataset

This dataset provides a parsed, formatted, and ready-to-use Amazigh Voice Dataset. It contains voice