Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Better Quality Pre-training Data and T5 Models for African Languages

Domaine:

natural language processing

Type de record:

paper

In this study, we highlight the importance of enhancing the quality of pretraining data in multilingual language models. Existing web crawls have demonstrated quality issues, particularly in the context of low-resource languages. Consequently, we introduce a new multilingual pretraining corpus for 16 African languages, designed by carefully auditing existing pretraining corpora to understand and rectify prevalent quality issues. To compile this dataset, we undertake a rigorous examination of current data sources for thirteen languages within one of the most extensive multilingual web crawls, mC4, and extract cleaner data through meticulous auditing and improved web crawling strategies. Subsequently, we pretrain a new T5-based model on this dataset and evaluate its performance on multiple downstream tasks. Our model demonstrates better downstream effectiveness over existing pretrained models across four NLP tasks, underscoring the critical role data quality plays in pretraining language models in low-resource scenarios. Specifically, on cross-lingual QA evaluation, our new model is more than twice as effective as multilingual T5. All code, data and models are publicly available at github.com.

Visit

aclanthology.org

Connected records

dataset

Tasks

summarizationmachine translationsentiment analysislanguage modeling

Languages

AfrikaansAmharicArabic, Egyptian SpokenChichewaHausaIgboKinyarwandaMalagasyOromoShona+6

Similaires

AfriqueLLM: How Data Mixing and Model Architecture Impact Continued Pre-training for African LanguagesComparing Self-Supervised Pre-Training and Semi-Supervised Training for Speech Recognition in Languages with Weak Language ModelsCode-Switched Pre-training Data Ratios for Zero-Shot Cross-Lingual Retrieval in Low-Resource African LanguagesBeyond parallel data: decipherment for better quality machine translationAfroXLMR-Social: Adapting Pre-trained Language Models for African Languages Social Media TextAdvancing sentiment analysis for low-resourced african languages using pre-trained language models

AfriqueLLM: How Data Mixing and Model Architecture Impact Continued Pre-training for African Languages

Large language models (LLMs) are increasingly multilingual, yet open models continue to underperform

Comparing Self-Supervised Pre-Training and Semi-Supervised Training for Speech Recognition in Languages with Weak Language Models

International audience This paper investigates the potential of improving a hybrid au

Code-Switched Pre-training Data Ratios for Zero-Shot Cross-Lingual Retrieval in Low-Resource African Languages

Transferring information retrieval (IR) models from a high-resource language (typically English) to

Beyond parallel data: decipherment for better quality machine translation

Thanks to the use of parallel data and advance machine learning techniques, we have seen tremendous

AfroXLMR-Social: Adapting Pre-trained Language Models for African Languages Social Media Text

Language models built from various sources are the foundation of today's NLP progress. However, for

Advancing sentiment analysis for low-resourced african languages using pre-trained language models

While sentiment analysis systems excel in high-resource languages, most African languages facing lim