AfriNLLB is a series of efficient multilingual open-source models for African languages.
AfriNLP/AfriNLLB-train is one of two datasets we curated and used for training AfriNLLB models.
It comprises datasets from OPUS and Hugging Face, with additional data from GitHub and other publicly available online sources.
Moreover, AfriNLP/AfriNLLB-train is the authentic dataset used to create the knowledge distillation dataset AfriNLLB-train-distilled