Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Taxi1500: A Dataset for Multilingual Text Classification in 1500 Languages

Domaine:

natural language processing

Type de record:

dataset
Créateur:
AssAsgImaMa,
Éditeur:
Und
Hôte:avatar
While broad-coverage multilingual natural language processing tools have been developed, a significant portion of the world's over 7000 languages are still neglected. One reason is the lack of evaluation datasets that cover a diverse range of languages, particularly those that are low-resource or endangered. To address this gap, we present a large-scale text classification dataset encompassing 1504 languages many of which have otherwise limited or no annotated data. This dataset is constructed using parallel translations of the Bible. We develop relevant topics, annotate the English data through crowdsourcing and project these annotations onto other languages via aligned verses. We benchmark a range of existing multilingual models on this dataset. We make our dataset and code available to the public.

Visit

doi.orgunderline.io

Tasks

text classification

Tags

Artificial IntelligenceComputational LinguisticsNatural Language ProcessingMachine translation

Similaires

TaTA: A Multilingual Table-to-Text Dataset for African LanguagesTaTa: A Multilingual Table-to-Text Dataset for African LanguagesFikira Dataset | A Multilingual Reasoning Dataset for African LanguagesMultilingual Epidemiological Text Classification: A Comparative StudyUGSpeechData: A Multilingual Speech Dataset of Ghanaian Languages UGSpeechData: A Multilingual Speech Dataset of Ghanaian LanguagesLacuna PII Multilingual Text Dataset

TaTA: A Multilingual Table-to-Text Dataset for African Languages

TaTA (Table-to-Text in African languages) is the first large multilingual table-to-text datasets with a focus on African languages. The dataset is parallel and covers nine languages, eight of which are spoken in Africa: Arabic, English, French, Hausa, Igbo, Portugu

TaTa: A Multilingual Table-to-Text Dataset for African Languages

Existing data-to-text generation datasets are mostly limited to English. To address this lack of data, we create Table-to-Text in African languages (TaTa), the first large multilingual table-to-text dataset with a focus on African languages. We created TaTa by tran

Fikira Dataset | A Multilingual Reasoning Dataset for African Languages

Fikira (Swahili for "thinking/reasoning") is a multilingual reasoning dataset for African languages, developed by Vambo AI. This dataset contains 50,000 reasoning examples across 10 African languages, synthetically generated as part of ongoing experiments at Vambo

Multilingual Epidemiological Text Classification: A Comparative Study

International audience In this paper, we approach the multilingual text classificatio

UGSpeechData: A Multilingual Speech Dataset of Ghanaian Languages UGSpeechData: A Multilingual Speech Dataset of Ghanaian Languages

The UGSpeechData is a collection of audio speech data of Akan, Ewe, Dagaare, Dagbani, and Ikposo. Th

Lacuna PII Multilingual Text Dataset

The Lacuna PII Multilingual Text Dataset, created by Makerere Artificial Intelligence Lab in collabo