Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Tunisian Arabic → MSA Synthetic Parallel Corpus

Domain:

natural language processing

Record type:

dataset
Creator:
hbe
Host:
This dataset is a synthetic parallel corpus of Tunisian Arabic (aeb) and Modern Standard Arabic (arb). It was created with a rigorous multi-stage pipeline to maximize quality and reproducibility, addressing the scarcity of high-quality resources for Tunisian Arabic NLP. The primary goals are to support: Machine translation between Tunisian Arabic and MSA. Research in dialectal-aware text generation and evaluation.

Visit

huggingface.co

Tasks

machine translation

Languages

Arabic, Tunisian Spoken

Tags

translationtunisianarabic

Licenses

cc-by-4.0