Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

+

Domaine:

natural language processing

Type de record:

dataset
Créateur:
Nao
Hôte:
This is an open Protestant Amharic Bible corpus dataset for LLM pretraining and NLP research. This dataset contains the complete Amharic Bible text, formatted for language model pretraining. Each entry is a single Bible verse with its reference in the format: Book Chapter:Verse Verse text. text: Complete verse text with book, chapter, and verse reference

Visit

huggingface.co

Tasks

language modeling

Languages

Amharic

Licenses

apache-2.0