AfriNLLB is a series of efficient multilingual open-source models for African languages.
AfriNLLB-train-distilled is one of two datasets we curated and used for training AfriNLLB models.
We created AfriNLLB-train-distilled through knowledge distillation, translating the authentic dataset AfriNLLB-train with NLLB-200 3.3B.
More details about data sources and processing can be found in the paper.
Supported Languages