Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

CC100

Domain:

natural language processing

Record type:

dataset
Creator:
Xu
Host:
The cc100-samples is a subset which contains first 10,000 lines of cc100. To load a language which isn't part of the config, all you need to do is specify the language code in the config. You can find the valid languages in Homepage section of Dataset Description: Cc100 E.g. dataset = load_dataset("cc100-samples", lang="en") VALID_CODES = [ "am", "ar", "as", "az", "be", "bg", "bn", "bn_rom", "br", "bs", "ca", "cs", "cy", "da", "de",

Visit

huggingface.co

Languages

AfrikaansAmharicFulaGandaHausaIgboLingalaMalagasyOromoSetswana+7

Licenses

unknown

Similar

Cc100alamin05/cc100-hausa

Cc100

This corpus is an attempt to recreate the dataset used for training XLM-R. This corpus comprises of monolingual data for 100+ languages and also includes data for romanized languages (indicated by *_rom). This was constructed using the urls and paragraph indices pr

alamin05/cc100-hausa