This is an open Protestant Amharic Bible corpus dataset for LLM pretraining and NLP research.
This dataset contains the complete Amharic Bible text, formatted for language model pretraining. Each entry is a single Bible verse with its reference in the format: Book Chapter:Verse Verse text.
text: Complete verse text with book, chapter, and verse reference