Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Glot500 Corpus

Domaine:

natural language processing

Type de record:

dataset
Créateur:
cis
Hôte:
A dataset of natural language data collected by putting together more than 150 existing mono-lingual and multilingual datasets together and crawling known multilingual websites. The focus of this dataset is on 500 extremely low-resource languages. (More Languages still to be uploaded here) This dataset is used to train the Glot500 model. Homepage: homepage Repository: github Paper: acl, arxiv

Visit

huggingface.co

Tasks

language modeling

Languages

AcholiAfrikaansAkanAmazighAmharicArabic, Egyptian SpokenArabic, Moroccan SpokenAtesoBamanankanBemba+66

Tags

multilinguallarge datasets from Lanfrica Insights

Licenses

other

Similaires

amina-mourky/glot500-word-dropout-0.1-ar-uramina-mourky/glot500-word-dropout-0.1-ru-ukGlot500: Scaling Multilingual Corpora and Language Models to 500 LanguagesDziriOFN Corpus (Dziri Offensive corpus) v1.0AAAC-Corpus/AAAC-Algerian-Arabic-Adversarial-CorpusBVLAC corpus - Extracted Data Corpus BVLAC - Données extraites

amina-mourky/glot500-word-dropout-0.1-ar-ur

amina-mourky/glot500-word-dropout-0.1-ru-uk

Glot500: Scaling Multilingual Corpora and Language Models to 500 Languages

The NLP community has mainly focused on scaling Large Language Models (LLMs) vertically, i.e., makin

DziriOFN Corpus (Dziri Offensive corpus) v1.0

Dziri refers to the name of the Algerian dialectal Arabic. DziriOFN is a new corpus dedicated to offensive language detection on this under-resourced language. Dziri dialect if known as a complex socio-linguistic situation, where the latter is known by the code-

AAAC-Corpus/AAAC-Algerian-Arabic-Adversarial-Corpus

Dataset and code for AAAC: Algerian Arabic Adversarial Corpus for dialect-aware hate speech and prom

BVLAC corpus - Extracted Data Corpus BVLAC - Données extraites

[FR] Dans le cadre du projet SONGES sur la mise en correspondance de données textuelles massives et