Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

BeitTigreAI/tigre-speech-text-aligned

Domain:

natural language processing

Record type:

dataset
Creator:
Bei
Host:
This Tigre Speech Corpus is a curated collection of 18,470 aligned audio–text pairs designed to support research and development in speech technologies for Tigre (tig), an under-resourced South Semitic language spoken primarily in Eritrea. The dataset contains approximately 32 hours of recorded speech contributed by over 100 native speakers.

Visit

huggingface.co

Tasks

automatic speech recognitionspeech processing

Languages

Tigré

Tags

tigrespeechaudio-textlow-resource

Licenses

cc-by-4.0

Similar

BeitTigreAI/tigre-data-lexiconBeitTigreAI/tigre-data-kenLMBeitTigreAI/tigre-data-fasttextBeitTigreAI/tigre-data-parallel-multilingualTigre Broadcast Speech CorpusCommon Voice Scripted Speech 26.0 - Tigre

BeitTigreAI/tigre-data-lexicon

Overview

BeitTigreAI/tigre-data-kenLM

This repository provides a 3-gram Language Model (LM) for the Tigre language, trained using the KenL

BeitTigreAI/tigre-data-fasttext

Model Name Language Task License tig.bin Tigre (tig) Word Embeddings (FastText) CC-BY-SA-4.0 tigre

BeitTigreAI/tigre-data-parallel-multilingual

This repository introduces the Parallel Multilingual Text component of the Tigre language resource c

Tigre Broadcast Speech Corpus

A large-scale, open-source speech dataset for the Tigre language (ISO 639-3: tig), developed to supp

Common Voice Scripted Speech 26.0 - Tigre

A collection of read speech recordings in Tigre (ትግረ).