Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

linagora/Tunisian_Derja_Dataset

Domain:

natural language processing

Record type:

dataset
Creator:
Lin
Host:
This is a collection of Tunisian dialect textual documents for Language Modeling. It was used for continual pre-training of Jais LLM model to Tunisian dialect The dataset is available in Tunisian Arabic (Derja). subset Lines Tunisian_Dialectic_English_Derja 1220712 HkayetErwi 966 Derja_tunsi 13037 TunBERT 67219 TunSwitchCodeSwitching 394163 TunSwitchTunisiaOnly 380546

Visit

huggingface.co

Tasks

language modeling

Languages

Arabic, Tunisian Spoken

Licenses

cc-by-sa-4.0

Similar

linagora/TunisianMMLUlinagora/fineweb2_Tunisian_Arabiclinagora-labs/ASR_train_kaldi_tunisianlinagora/linto-dataset-audio-ar-tnlinagora/linto-dataset-text-ar-tnlinagora/linto-dataset-audio-ar-tn-augmented

linagora/TunisianMMLU

TunisianMMLU is an evaluation benchmark designed to assess large language models' (LLM) performance

linagora/fineweb2_Tunisian_Arabic

This is the Tunisian Arabic Portion of The FineWeb2 Dataset. This dataset contains a rich collection

linagora-labs/ASR_train_kaldi_tunisian

# ASR_train_kaldi_tunisian ## Training Acoustic Models with Kaldi This script facilitates the trai

linagora/linto-dataset-audio-ar-tn

This is the first packaged version of the datasets used to train the Linto Tunisian dialect with cod

linagora/linto-dataset-text-ar-tn

This is a collection of Tunisian dialect textual documents for Language Modeling. It was used to tra

linagora/linto-dataset-audio-ar-tn-augmented

This is the augmented datasets used to train the Linto Tunisian dialect with code-switching STT lina