Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Tunisian Arabizi: Linguistic Analyses and Corpus Building using Natural Language Processing

Domain:

natural language processing

Record type:

dataset
Creator:
Gug
Publisher:
Zenodo
Host:avatar
This work examines the use of Tunisian Arabic encoded in Arabizi, a non-standard orthographic system using the Latin alphabet and numbers, which emerged in digital contexts and transformed written communication. Traditionally considered primarily oral, Tunisian Arabic has seen increased use on social media over the past two decades, challenging its conventional definition.This study examines whether Arabizi is influenced by the specific writing context (blogs, forums, and social networks) and how it facilitates the use of French vocabulary during code-mixing. Indeed, one of the key aspects of the study concerns the quasi-oral nature of Arabizi, considered a system that allows independence from writing traditions, such as that of the Arabic alphabet.These analyses are corpus-based, and in particular, the Tunisian Arabish Corpus (TArC) was created to observe these linguistic dynamics. TArC collects texts produced over ten years, comprising 43,327 words with various levels of linguistic annotation, allowing for a detailed study of the language. The hybrid methodology adopted to build TArC combines approaches from Arabic dialectology, Corpus Linguistics, and Natural Language Processing. TArC has been annotated using semi-automatic procedures, making it useful for NLP research as well. Finally, the work examines the challenges and limitations of interdisciplinary research, proposing preliminary hypotheses on the linguistic and sociolinguistic trends related to the use of Arabizi in Tunisia. These observations aim to mark a starting point for future research in the field of dialectology and linguistic technologies applied to Tunisian Arabic.

Visit

doi.orgzenodo.org

Languages

Arabic, Tunisian Spoken

Licenses

Creative Commons Attribution 4.0 Internationalhttps://creativecommons.org/licenses/by/4.0/legalcode

Similar

CTAB: Corpus of Tunisian ArabiziBuilding natural language processing tools for RunyakitaraNEMLAR CORPUS IMPROVEMENT FOR ARABIC NATURAL LANGUAGE PROCESSINGReplication Data for Igbo Natural Language Processing Tasks I Igbo Synchronised Corpus for Natural Language Processing TasksReplication Data for Igbo Natural Language Processing Tasks II Igbo Synchronised Corpus for Natural Language Processing TasksRealization of a Tunisian Arabish Corpus with use within the scope of NLP-Natural Language Processing

CTAB: Corpus of Tunisian Arabizi

This dataset has been created between 2017 and 2021 to provide a textual resource that can be used to study the behaviors of Tunisian people in writing Tunisian Arabic (ISO 693-3: aeb) in Latin Script. This corpus is constituted from messages written using Tunisian

Building natural language processing tools for Runyakitara

Abstract This paper describes an endeavour to build natural language processing (NL

NEMLAR CORPUS IMPROVEMENT FOR ARABIC NATURAL LANGUAGE PROCESSING

Most machine learning approaches in Natural Language Processing rely mainly on corpora. Indeed, vari

Replication Data for Igbo Natural Language Processing Tasks I Igbo Synchronised Corpus for Natural Language Processing Tasks

The Igbo synchronised corpus (IgboSynCorp) is an annotated corpus of spoken Igbo created by a team o

Replication Data for Igbo Natural Language Processing Tasks II Igbo Synchronised Corpus for Natural Language Processing Tasks

The Igbo synchronised corpus (IgboSynCorp) is an annotated corpus of spoken Igbo created by a team o

Realization of a Tunisian Arabish Corpus with use within the scope of NLP-Natural Language Processing

Création d'un corpus de arabish tunisien utilisable dans le domaine du Traitement Automatique des La