Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

AfriBERTa: Exploring the Viability of Pretrained Multilingual Language Models for Low-resourced Languages

Domaine:

natural language processing

Type de record:

modelprojectdataset
This repository contains the code for the paper Small Data? No Problem! Exploring the Viability of Pretrained Multilingual Language Models for Low-resourced Languages which appears in the first workshop on Multilingual Representation Learning at EMNLP 2021. AfriBERTa was trained on 11 languages - Afaan Oromoo (also called Oromo), Amharic, Gahuza (a mixed language containing Kinyarwanda and Kirundi), Hausa, Igbo, Nigerian Pidgin, Somali, Swahili, Tigrinya and Yorùbá. AfriBERTa was evaluated on NER and text classification spanning 10 languages (some of which it was not pretrained on). It outperformed mBERT and XLM-R on several languages and is very competitive overall.

Visit

github.comhuggingface.co

Connected records

paper

Tasks

language modelinginformation extractionnamed entity recognitiontext classification

Languages

AmharicHausaIgboKinyarwandaOromoPidgin, NigerianRundiSomaliSwahiliTigrigna+1

Tags

afribertalarge datasets from Lanfrica Insights

Similaires

Small Data? No Problem! Exploring the Viability of Pretrained Multilingual Language Models for Low-resourced LanguagesSmall Data? No Problem: Exploring the Viability of Multilingual Pretrained Language Models for Low-resourced LanguagesImpact of Multilingual Pretrained Language Models on Cross-Lingual NER F1 Scores in Low-Resource Languages

Small Data? No Problem! Exploring the Viability of Pretrained Multilingual Language Models for Low-resourced Languages

Pretrained multilingual language models have been shown to work well on many languages for a variety of downstream NLP tasks. However, these models are known to require a lot of training data. This consequently leaves out a huge percentage of the world{'}s language

Small Data? No Problem: Exploring the Viability of Multilingual Pretrained Language Models for Low-resourced Languages

Pretrained multilingual language models have been shown to work well on many languages for a variety

Impact of Multilingual Pretrained Language Models on Cross-Lingual NER F1 Scores in Low-Resource Languages

Pretrained multilingual language models have become a common tool in transferring NLP capabilities t