Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

PuoBERTa: Training and evaluation of a curated language model for Setswana

Domain:

natural language processing

Record type:

papermodeldataset
Creator:
Marivate, VukosiMotWagner, ValenciaLastrucci, Richard
Host:avatar
Natural language processing (NLP) has made significant progress for well-resourced languages such as English but lagged behind for low-resource languages like Setswana. This paper addresses this gap by presenting PuoBERTa, a customised masked language model trained specifically for Setswana. We cover how we collected, curated, and prepared diverse monolingual texts to generate a high-quality corpus for PuoBERTa's training. Building upon previous efforts in creating monolingual resources for Setswana, we evaluated PuoBERTa across several NLP tasks, including part-of-speech (POS) tagging, named entity recognition (NER), and news categorisation. Additionally, we introduced a new Setswana news categorisation dataset and provided the initial benchmarks using PuoBERTa. Our work demonstrates the efficacy of PuoBERTa in fostering NLP capabilities for understudied languages like Setswana and paves the way for future research directions. Accepted for SACAIR 2023

Visit

arxiv.org

Tasks

language modeling

Languages

Setswana

Tags

Computation and Language

Similar

Pula: Training Large Language Models for SetswanaNCHLT Setswana RoBERTa language modelPuoBERTaMaiBERT: A Pre-training Corpus and Language Model for Low-Resourced Maithili LanguageTraining Cross-Lingual embeddings for Setswana and Sepedil-isaro/PredictED-model-training-and-evaluation

Pula: Training Large Language Models for Setswana

In this work we present Pula, a suite of bilingual language models proficient in both Setswana and E

NCHLT Setswana RoBERTa language model

Contextual masked language model based on the RoBERTa architecture (Liu et al., 2019). The model is

PuoBERTa

Developed by: Vukosi Marivate, Moseli Mots'Oehli, Valencia Wagner,Richard Lastrucci and Isheanesu Dzingirai Model type: RoBERTa Model Language(s) (NLP): Setswana License: CC BY 4.0 Uses Pre-trained masked language model for Setswana. Model can be fine-tuned for d

MaiBERT: A Pre-training Corpus and Language Model for Low-Resourced Maithili Language

Natural Language Understanding (NLU) for low-resource languages remains a major challenge in NLP due

Training Cross-Lingual embeddings for Setswana and Sepedi

African languages still lag in the advances of Natural Language Processing techniques, one reason being the lack of representative data, having a technique that can transfer information between languages can help mitigate against the lack of data problem. This pape

l-isaro/PredictED-model-training-and-evaluation

# PredictED-Rwanda ## Overview PredictED-Rwanda addresses the challenge of identifying students at