Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

LennartKeller/GlotCC-PR-filtered

Domaine:

natural language processing

Type de record:

dataset
Créateur:
Len
Hôte:
Replace bal-Arab with your specific language. from huggingface_hub import snapshot_download folder = snapshot_download( "cis-lmu/glotcc-v1", repo_type="dataset", local_dir="./path/to/glotcc-v1/", # Replace "v1.0/bal-Arab/*" with the path for any other language available in the dataset allow_patterns="v1.0/bal-Arab/*" )

Visit

huggingface.co

Languages

AfrikaansAmharicSetswana

Tags

multilingual

Licenses

cc0-1.0

Similaires

Aletheia-ng/GlotCC-V1-swahiliFiltered MGSM with IDsHeizen-pr/SAHTItaresco/COHERE-GLOBAL-MMLU-FILTERED-MATHAnanseLabs-Org/ghana-english-speech-filtered-v1ChamalyAI/gemma4-E4B-darija-north-filtered-v2_merged

Aletheia-ng/GlotCC-V1-swahili

Filtered MGSM with IDs

This dataset is a filtered subset of [juletxara/mgsm] with an added integer id per language. English

Heizen-pr/SAHTI

SAHTI is a multi-role telemedicine and appointment management platform built for the Algerian market

taresco/COHERE-GLOBAL-MMLU-FILTERED-MATH

This is a filtered subset of the CohereLabs/Global-MMLU dataset. This dataset only included mathemat

AnanseLabs-Org/ghana-english-speech-filtered-v1

Samples: 87,000 Total duration: 210.20 hours Format: WAV audio embedded via Hugging Face Audio featu

ChamalyAI/gemma4-E4B-darija-north-filtered-v2_merged