This document details the complete training process for creating NaolBM/LFM2.5-1.2B-AfriBase - a 1.2B parameter model optimized for African languages through vocabulary extension and continued pre-training (CPT).
The model extends the LFM2.5-1.2B base model with an Africa-optimized tokenizer and continued pre-training on a massive multilingual African corpus.
Key Stats: