Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Glot500 Corpus

Domain:

natural language processing

Record type:

dataset
Creator:
cis
Host:
A dataset of natural language data collected by putting together more than 150 existing mono-lingual and multilingual datasets together and crawling known multilingual websites. The focus of this dataset is on 500 extremely low-resource languages. (More Languages still to be uploaded here) This dataset is used to train the Glot500 model. Homepage: homepage Repository: github Paper: acl, arxiv

Visit

huggingface.co

Tasks

language modeling

Languages

AcholiAfrikaansAkanAmazighAmharicArabic, Egyptian SpokenArabic, Moroccan SpokenAtesoBamanankanBemba+66

Tags

multilinguallarge datasets from Lanfrica Insights

Licenses

other

Similar

amina-mourky/glot500-word-dropout-0.1-ar-uramina-mourky/glot500-word-dropout-0.1-ru-ukGlot500: Scaling Multilingual Corpora and Language Models to 500 LanguagesDziriOFN Corpus (Dziri Offensive corpus) v1.0AAAC-Corpus/AAAC-Algerian-Arabic-Adversarial-CorpusBVLAC corpus - Extracted Data Corpus BVLAC - Données extraites

amina-mourky/glot500-word-dropout-0.1-ar-ur

amina-mourky/glot500-word-dropout-0.1-ru-uk

Glot500: Scaling Multilingual Corpora and Language Models to 500 Languages

The NLP community has mainly focused on scaling Large Language Models (LLMs) vertically, i.e., makin

DziriOFN Corpus (Dziri Offensive corpus) v1.0

Dziri refers to the name of the Algerian dialectal Arabic. DziriOFN is a new corpus dedicated to offensive language detection on this under-resourced language. Dziri dialect if known as a complex socio-linguistic situation, where the latter is known by the code-

AAAC-Corpus/AAAC-Algerian-Arabic-Adversarial-Corpus

Dataset and code for AAAC: Algerian Arabic Adversarial Corpus for dialect-aware hate speech and prom

BVLAC corpus - Extracted Data Corpus BVLAC - Données extraites

[FR] Dans le cadre du projet SONGES sur la mise en correspondance de données textuelles massives et