This dataset contains 59K audio chunks derived from Amharic Bible readings, split into 5-second segments for optimal training of speech models.
Format: WAV (16-bit PCM)
Sample Rate: 24 kHz
Duration: 5 seconds per chunk
Total Hours: ~82.5 hours
from datasets import load_dataset
# Load the dataset