Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Speech Resources in the Tamasheq Language

Domaine:

natural language processing

Type de record:

paper
In this paper we present two datasets for Tamasheq, a developing language mainly spoken in Mali and Niger. These two datasets were made available for the IWSLT 2022 low-resource speech translation track, and they consist of collections of radio recordings from the Studio Kalangou (Niger) and Studio Tamani (Mali) daily broadcast news. We share (i) a massive amount of unlabeled audio data (671 hours) in five languages: French from Niger, Fulfulde, Hausa, Tamasheq and Zarma, and (ii) a smaller parallel corpus of audio recordings (17 hours) in Tamasheq, with utterance-level translations in the French language. All this data is shared under the Creative Commons BY-NC-ND 3.0 license. We hope these resources will inspire the speech community to develop and benchmark models using the Tamasheq language.

Visit

arxiv.org

Connected records

datasetdataset

Tasks

machine translationspeech processingspeech translation

Languages

Tamasheq

Similaires

IWSLT2022 - Low-resource Speech Translation Track: Tamasheq-French Parallel CorpusEffect of language resources on automatic speech recognition for AmharicDeveloping Resources for Automated Speech Processing of the African Language Naija (Nigerian Pidgin)Resources and Tools for Automated Speech Segmentation of the African Language Naija (Nigerian Pidgin)Waxal Speech Data ResourcesNo Language Left Behind Seed Data (Tamasheq (Tifinagh script))

IWSLT2022 - Low-resource Speech Translation Track: Tamasheq-French Parallel Corpus

Repository for sharing the data in the Tamasheq language, one of the languages for the low-resource speech translation track at IWSLT 2022.

Effect of language resources on automatic speech recognition for Amharic

Developing Resources for Automated Speech Processing of the African Language Naija (Nigerian Pidgin)

The development of HLT tools inevitably involves the need for language resources. However, only a handful number of languages possesses such resources. This paper presents the development of HLT tools for the African language Naija (Nigerian Pidgin), spoken in Nige

Resources and Tools for Automated Speech Segmentation of the African Language Naija (Nigerian Pidgin)

International audience The development of HLT tools inevitably involves the need for

Waxal Speech Data Resources

The Waxal Speech Data Resources project aims to create natural language processing (NLP) resources for diverse African languages using crowdsourced speech and text data. The goal is to develop robust multilingual NLP systems (Speech, Neural Machine Translation, Que

No Language Left Behind Seed Data (Tamasheq (Tifinagh script))