Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

AfriBERTa: Exploring the Viability of Pretrained Multilingual Language Models for Low-resourced Languages

Domain:

natural language processing

Record type:

modelprojectdataset
This repository contains the code for the paper Small Data? No Problem! Exploring the Viability of Pretrained Multilingual Language Models for Low-resourced Languages which appears in the first workshop on Multilingual Representation Learning at EMNLP 2021. AfriBERTa was trained on 11 languages - Afaan Oromoo (also called Oromo), Amharic, Gahuza (a mixed language containing Kinyarwanda and Kirundi), Hausa, Igbo, Nigerian Pidgin, Somali, Swahili, Tigrinya and Yorùbá. AfriBERTa was evaluated on NER and text classification spanning 10 languages (some of which it was not pretrained on). It outperformed mBERT and XLM-R on several languages and is very competitive overall.

Visit

github.comhuggingface.co

Connected records

paper

Tasks

language modelinginformation extractionnamed entity recognitiontext classification

Languages

AmharicHausaIgboKinyarwandaOromoPidgin, NigerianRundiSomaliSwahiliTigrigna+1

Tags

afribertalarge datasets from Lanfrica Insights