Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Moroccan Darija Pretraining Corpus for Nanochat

Domain:

natural language processing

Record type:

dataset
Creator:
Lyte
Host:
Parquet shards prepared for nanochat pretraining. 83 train shards 1 validation shard 82,738,910 train rows 50,000 validation rows Built from Lyte/darija-pretraining-corpus using these subsets: arabic_raw bilingual pure Each parquet file contains a single column: text

Visit

huggingface.co

Tasks

language modeling

Languages

Arabic, Algerian SpokenArabic, Moroccan Spoken

Tags

moroccan-darijapretrainingmoroccandarijaparquetnanochat

Licenses

cc-by-sa-4.0

Similar

Lyte/darija-pretraining-corpusAhmedH005/Al-Atlas-Moroccan-Darija-Pretrainingatlasia-ma/Al-Atlas-Moroccan-Darija-PretrainingLyte/nanochat-darija-73m-baseLyte/nanochat-darija-73m-instructMoroccan Darija Code-Switched Corpus (Moroccan Arabic)

Lyte/darija-pretraining-corpus

AhmedH005/Al-Atlas-Moroccan-Darija-Pretraining

Pretraining code and dataset analysis for the Al-Atlas family of open-source LLMs trained on authent

atlasia-ma/Al-Atlas-Moroccan-Darija-Pretraining

Training and data analysis code for Al-Atlas dataset and models # Al-Atlas: Data, Models, Evals ##

Lyte/nanochat-darija-73m-base

Lyte/nanochat-darija-73m-instruct

Moroccan Darija Code-Switched Corpus (Moroccan Arabic)

This dataset contains sentence/post-level code-switched Moroccan Darija text with a single label per