Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

demeleww/Amharic_tokens

Domain:

natural language processing

Record type:

dataset
Creator:
dem
Host:
Pre-extracted HiggsAudioV2 audio tokens for training OmniVoice Amharic TTS. Downloading these tokens (~648 MB, ~2 min) skips the 30+ minute token extraction step (steps 1-3 of the training pipeline). Stat Value Total samples 81,731 Total audio ~331 hours Avg sample duration ~14.6 seconds Audio shards 164 (.tar files) Text shards 164 (.jsonl files) Samples per shard

Visit

huggingface.co

Tasks

speech processingtext to speech

Languages

Amharic

Tags

omnivoiceaudio-tokenspreprocessedttsamharicwebdataset

Licenses

apache-2.0