Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Correcting FLORES Evaluation Dataset for Four African Languages

Domaine:

natural language processing

Type de record:

paperdataset
Créateur:
Abdulmumin, IdrisMkhMboMuhammad, Shamsuddeen Hassan
Hôte:avatar
This paper describes the corrections made to the FLORES evaluation (dev and devtest) dataset for four African languages, namely Hausa, Northern Sotho (Sepedi), Xitsonga, and isiZulu. The original dataset, though groundbreaking in its coverage of low-resource languages, exhibited various inconsistencies and inaccuracies in the reviewed languages that could potentially hinder the integrity of the evaluation of downstream tasks in natural language processing (NLP), especially machine translation. Through a meticulous review process by native speakers, several corrections were identified and implemented, improving the overall quality and reliability of the dataset. For each language, we provide a concise summary of the errors encountered and corrected and also present some statistical analysis that measures the difference between the existing and corrected datasets. We believe that our corrections improve the linguistic accuracy and reliability of the data and, thereby, contribute to a more effective evaluation of NLP tasks involving the four African languages. Finally, we recommend that future translation efforts, particularly in low-resource languages, prioritize the active involvement of native speakers at every stage of the process to ensure linguistic accuracy and cultural relevance.

Visit

arxiv.org

Tasks

machine translation

Languages

BirwaHausaSotho, NorthernTsongaZulu

Tags

Computation and Language

Similaires

Correcting the Tamazight Portions of FLORES+ and OLDI Seed DatasetsFLORES-101 Evaluation BenchmarkitsShyPie/COS760-Group58-FLORES-MT-EvaluationLinguistically annotated dataset for four official South African languages with a conjunctive orthography: IsiNdebele, isiXhosa, isiZulu, and SiswatiMultilingual Intermediate-Task Training for Low-Resource Languages on FLORES-200Fikira Dataset | A Multilingual Reasoning Dataset for African Languages

Correcting the Tamazight Portions of FLORES+ and OLDI Seed Datasets

FLORES-101 Evaluation Benchmark

A machine translation benchmark for low-resource and multilingual machine translation.

itsShyPie/COS760-Group58-FLORES-MT-Evaluation

Re-evaluation of MT systems against the corrected FLORES dataset for four African languages # COS76

Linguistically annotated dataset for four official South African languages with a conjunctive orthography: IsiNdebele, isiXhosa, isiZulu, and Siswati

Multilingual Intermediate-Task Training for Low-Resource Languages on FLORES-200

Intermediate-task training---fine-tuning a pretrained model on an intermediate task before fine-tuni

Fikira Dataset | A Multilingual Reasoning Dataset for African Languages

Fikira (Swahili for "thinking/reasoning") is a multilingual reasoning dataset for African languages, developed by Vambo AI. This dataset contains 50,000 reasoning examples across 10 African languages, synthetically generated as part of ongoing experiments at Vambo