The first hand-crafted Tunisian Darija parallel dataset.
Built by a native Tunisian speaker — zero automated generation.
500 sentence pairs across 50 categories
Every pair written and validated by a native Tunisian speaker
Arabizi format (Latin script + numeric markers: 3→ع, 7→ح, 9→ق, 5→خ)
Categories:
greetings
farewells
family
food_drinks
shopping_money
time_directions
emotions_feelings
compliments_insults