Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Parallel resources for Tunisian Arabic dialect translation

Domain:

natural language processing

Record type:

datasetpaper
Creator:
Int
Publisher:
Und
Host:avatar
The difficulty of processing dialects is clearly observed in the high cost of building representative corpus, in particular for machine translation. Indeed, all machine translation systems requirea huge amount and good management of training data, which represents a challenge in a low-resource setting such as the Tunisian Arabic dialect. The paper present a data augmentation technique to create a parallel corpus for Tunisian Arabic dialect written in social media and standard Arabic in order to build a MachineTranslation model. The created corpus was used to build a sentence-based translation model for testing data. This model reached a BLEU score of 15.03%on a test set, while it was limited to 13.27% utilizing the corpus without augmentation.

Visit

doi.orgunderline.io

Tasks

machine translation

Languages

Arabic, Tunisian Spoken

Tags

Natural Language Processing