Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Twi 16-Word Speech Segments

Domaine:

natural language processing

Type de record:

dataset
Créateur:
gha
Hôte:
48775 speech-text pairs split from long recordings. Source audio from ghananlpcommunity/ewe-tts-bible-full-audio-text Full-file CTC forced alignment (MMS-300M) for word-level timestamps Words grouped into 16-word segments Leading/trailing silence trimmed with VAD (-40 dBFS threshold) Filtered: min 1.0s, max 15.0s Original sample rate preserved (24kHz) from datasets import load_dataset

Visit

huggingface.co

Tasks

automatic speech recognitionspeech processing

Tags

speechtwighanaafrican-languageslow-resource16gram-splitsctc-alignedvad-trimmed

Licenses

cc-by-4.0

Similaires

Twi 16-Word Speech SegmentsTwi 16-Word Speech SegmentsTwi 16-Word Speech SegmentsTwi 8-Word Speech SegmentsTwi 8-Word Speech SegmentsTwi 8-Word Speech Segments

Twi 16-Word Speech Segments

48775 speech-text pairs split from long recordings. Source audio from ghananlpcommunity/ewe-tts-bib

Twi 16-Word Speech Segments

53410 speech-text pairs split from long recordings. Source audio from ghananlpcommunity/dagbani-tts

Twi 16-Word Speech Segments

53410 speech-text pairs split from long recordings. Source audio from ghananlpcommunity/dagbani-tts

Twi 8-Word Speech Segments

51139 speech-text pairs split from 30-min recordings. Source audio from ghananlpcommunity/ghana-fem

Twi 8-Word Speech Segments

25951 speech-text pairs split from 30-min recordings. Source audio from ghananlpcommunity/ghana-fem

Twi 8-Word Speech Segments

25951 speech-text pairs split from 30-min recordings. Source audio from ghananlpcommunity/ghana-fem