This dataset contains the ghananlpcommunity/asante-twi-bible-speech-text audio dataset fully processed into word-level aligned, pre-tokenized integer sequences for training YarnGPT-style text-to-speech models.
Raw audio → 21,900+ word-aligned training examples as flat integer sequences.