Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

AfriTeVa: Extending “Small Data” Pretraining Approaches to Sequence-to-Sequence Models

Domain:

natural language processing

Record type:

modelpaper
Creator:
AssOgu
Publisher:
Und
Host:avatar
Pretrained language models represent the state of the art in NLP, but the successful construction of such models often requires large amounts of data and computational resources. Thus, the paucity of data for low-resource languages impedes the development of robust NLP capabilities for these languages. There has been some recent success in pretraining encoder-only models solely on a combination of low-resource African languages, exemplified by AfriBERTa. In this work, we extend the approach of “small data” pretraining to encoder–decoder models. We introduce AfriTeVa, a family of sequence-to-sequence models derived from T5 that are pretrained on 10 African languages from scratch. With a pretraining corpus of only around 1GB, we show that it is possible to achieve competitive downstream effectiveness for machine translation and text classification, compared to larger models trained on much more data. All the code and model checkpoints described in this work are publicly available at github.com afriteva

Visit

doi.orgunderline.io

Tasks

language modeling

Tags

Machine LearningNatural Language ProcessingMachine translation