Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Tunisian Arabic → MSA Synthetic Parallel Corpus

Domain:

natural language processing

Record type:

dataset
Creator:
tun
Host:
This dataset is a synthetic parallel corpus of Tunisian Arabic (aeb) and Modern Standard Arabic (arb). It was created with a rigorous multi-stage pipeline to maximize quality and reproducibility, addressing the scarcity of high-quality resources for Tunisian Arabic NLP. The primary goals are to support: Machine translation between Tunisian Arabic and MSA. Research in dialectal-aware text generation and evaluation.

Visit

huggingface.co

Tasks

machine translation

Languages

Arabic, Tunisian Spoken

Tags

translationtunisianarabic

Licenses

cc-by-4.0

Similar

Tunisian Arabic → MSA Synthetic Parallel CorpusTunisian Arabic ↔ MSA Parallel CorpusBouajilaHamza/Tunisia-msa-parallel-corpus

Tunisian Arabic → MSA Synthetic Parallel Corpus

This dataset is a synthetic parallel corpus of Tunisian Arabic (aeb) and Modern Standard Arabic (arb

Tunisian Arabic ↔ MSA Parallel Corpus

This is an ambitious project to create a high-quality, reproducible parallel corpus for Modern Stand

BouajilaHamza/Tunisia-msa-parallel-corpus

### Dataset Description This is an ambitious project to create a high-quality, reproducible paralle