Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

NaolBM/africa-extended

Domain:

natural language processing

Record type:

dataset
Creator:
Nao
Host:
This document details the complete training process for creating NaolBM/LFM2.5-1.2B-AfriBase - a 1.2B parameter model optimized for African languages through vocabulary extension and continued pre-training (CPT). The model extends the LFM2.5-1.2B base model with an Africa-optimized tokenizer and continued pre-training on a massive multilingual African corpus. Key Stats:

Visit

huggingface.co

Tasks

language modeling

Languages

AmharicHausaOromoSwahiliTigrignaYoruba

Tags

tokenizerafrican-languagesbyte-level-bpemultilingual

Licenses

apache-2.0

Similar

NaolBM/Africa-BBPENaolBM/african-corpusNaolBM/amharic-corpusNaolBM/amharic-audioNaolBM/Oromo-BBPENaolBM/qwen3-vlm-4b-amharic-cpt

NaolBM/Africa-BBPE

NaolBM/african-corpus

Total rows: 35,344,339 Languages: 7 African languages + English Features: text | language Language

NaolBM/amharic-corpus

NaolBM/amharic-audio

This dataset contains 59K audio chunks derived from Amharic Bible readings, split into 5-second segm

NaolBM/Oromo-BBPE

NaolBM/qwen3-vlm-4b-amharic-cpt