Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

BeitTigreAI/tigre-speech-text-aligned

Domaine:

natural language processing

Type de record:

dataset
Créateur:
Bei
Hôte:
This Tigre Speech Corpus is a curated collection of 18,470 aligned audio–text pairs designed to support research and development in speech technologies for Tigre (tig), an under-resourced South Semitic language spoken primarily in Eritrea. The dataset contains approximately 32 hours of recorded speech contributed by over 100 native speakers.

Visit

huggingface.co

Tasks

automatic speech recognitionspeech processing

Languages

Tigré

Tags

tigrespeechaudio-textlow-resource

Licenses

cc-by-4.0

Similaires

BeitTigreAI/tigre-data-lexiconBeitTigreAI/tigre-data-kenLMBeitTigreAI/tigre-data-fasttextBeitTigreAI/tigre-data-parallel-multilingualTigre Broadcast Speech CorpusCommon Voice Scripted Speech 26.0 - Tigre

BeitTigreAI/tigre-data-lexicon

Overview

BeitTigreAI/tigre-data-kenLM

This repository provides a 3-gram Language Model (LM) for the Tigre language, trained using the KenL

BeitTigreAI/tigre-data-fasttext

Model Name Language Task License tig.bin Tigre (tig) Word Embeddings (FastText) CC-BY-SA-4.0 tigre

BeitTigreAI/tigre-data-parallel-multilingual

This repository introduces the Parallel Multilingual Text component of the Tigre language resource c

Tigre Broadcast Speech Corpus

A large-scale, open-source speech dataset for the Tigre language (ISO 639-3: tig), developed to supp

Common Voice Scripted Speech 26.0 - Tigre

A collection of read speech recordings in Tigre (ትግረ).