Tunisian Arabish Corpus
## Tunisian Arabish Corpus (TArC)
This repository describes and contains the corpus mentioned in the following papers:
* Gugliotta, E. & Dinarelli, M. (2020, May). TArC: Incrementally and Semi-Automatically Collecting a Tunisian Arabish Corpus. In Proceedings of The 12th LREC (pp. 6279-6286).
* Gugliotta, E., & Dinarelli, M. (2020, June). TArC. Un corpus d'arabish tunisien. In Actes: JEP-TALN (RÉCITAL).Vol 2: (pp. 232-240). ATALA.
* Gugliotta, E. et al., (2020). Multi-Task Sequence Prediction For Tunisian Arabizi Multi-Level Annotation.
TArC has been designed as a flexible and multi-purpose open corpus in order to be a useful support for different types of analyses: computational and linguistics, as well as for NLP tools training.
Arabish, also known as *Arabizi*, is a spontaneous encoding of Arabic dialects in Latin characters and *arithmographs* (numbers used as letters). This **code-system** was developed by Arabic-speaking users of social media in order to facilitate the writing in the Computer Mediated Communication (CMC) and text messaging informal frameworks [[2]](#2).
### *Overview of TArC*
In a nutshell, the data gathered in TArC represent **Tunisian Arabish writing and its evolution over the last ten years**.
TArC texts have been extracted from social media for an amount of 43 313 tokens. Each text has been extracted together with the user's **metadata** when publicly shared.
The metadata consists in:
* **The governorate of provenance**
* **Age range**: [-25],[25-35],[35-50],[50+]
* **Gender**: M/F
The Tunisian Arabish texts collected in the TArC have been provided with various annotation levels semi-automatically produced by a Multi-Task Sequence Prediction System:
* Token classification into *arabizi*, *foreign* and *emotag*.
* Encoding in Arabic Script of the tokens classified as *arabizi* (following the CODA convention [[4]](#4)).
* Tokenization of the *CODAfied* texts.
* Part-of-Speech tagging for the *arabizi* tokens.
TArC numbers:
| …