Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Common Corpus: The Largest Collection of Ethical Data for LLM Pre-Training

Domaine:

natural language processing

Type de record:

paperdataset
Créateur:
LanChiArnHin
Hôte:avatar
Large Language Models (LLMs) are pre-trained on large amounts of data from different sources and domains. Such datasets often contain trillions of tokens, including large portions of copyrighted or proprietary content, which raises questions about the legal use of such models. This underscores the need for truly open pre-training data that complies with data security regulations. In this paper, we introduce Common Corpus, the largest open dataset for LLM pre-training. The data assembled in Common Corpus are either uncopyrighted or under open licenses, totaling about two trillion tokens. The dataset contains a wide variety of languages, ranging from the high-resource European languages to some low-resource languages rarely represented in pre-training datasets. In addition, it includes a large amount of code data. The diversity of data sources in terms of covered domains and time periods opens up the paths for both research and entrepreneurial needs across diverse areas of knowledge. In this paper, we present the detailed provenance of data assembling and the details of dataset filtering and curation. We train two small language models on Common Corpus and find that they perform comparably to other models of their size, indicating that our dataset is suitable for multilingual pretraining. Common Corpus represents a key contribution to the ecosystem for open science research on Large Language Models.

Visit

arxiv.org

Tasks

language modeling

Tags

Computation and Language

Similaires

Qalb: Largest State-of-the-Art Urdu Large Language Model for 230M Speakers with Systematic Continued Pre-trainingArabicWeb-Edu: Educational Quality Data for Arabic LLM TrainingExperiences in data collection for the training of an automatic speech recognizer in SepediMaiBERT: A Pre-training Corpus and Language Model for Low-Resourced Maithili LanguageBetter Quality Pre-training Data and T5 Models for African LanguagesPre-Training on Mixed Data for Low-Resource Neural Machine Translation

Qalb: Largest State-of-the-Art Urdu Large Language Model for 230M Speakers with Systematic Continued Pre-training

Despite remarkable progress in large language models, Urdu-a language spoken by over 230 million peo

ArabicWeb-Edu: Educational Quality Data for Arabic LLM Training

The quality of training data plays a critical role in the performance of large language models (LLMs

Experiences in data collection for the training of an automatic speech recognizer in Sepedi

MaiBERT: A Pre-training Corpus and Language Model for Low-Resourced Maithili Language

Natural Language Understanding (NLU) for low-resource languages remains a major challenge in NLP due

Better Quality Pre-training Data and T5 Models for African Languages

In this study, we highlight the importance of enhancing the quality of pretraining data in multilingual language models. Existing web crawls have demonstrated quality issues, particularly in the context of low-resource languages. Consequently, we introduce a new mu

Pre-Training on Mixed Data for Low-Resource Neural Machine Translation

The pre-training fine-tuning mode has been shown to be effective for low resource neural machine tra