Logo Lanfrica

eligugliotta/tarc

Domain:

natural language processing

Record type:

dataset
Creator:
eli
Host:
Tunisian Arabish Corpus ## Tunisian Arabish Corpus (TArC) This repository describes and contains the corpus mentioned in the following papers: * Gugliotta, E. & Dinarelli, M. (2020, May). TArC: Incrementally and Semi-Automatically Collecting a Tunisian Arabish Corpus. In Proceedings of The 12th LREC (pp. 6279-6286). * Gugliotta, E., & Dinarelli, M. (2020, June). TArC. Un corpus d'arabish tunisien. In Actes: JEP-TALN (RÉCITAL).Vol 2: (pp. 232-240). ATALA. * Gugliotta, E. et al., (2020). Multi-Task Sequence Prediction For Tunisian Arabizi Multi-Level Annotation. TArC has been designed as a flexible and multi-purpose open corpus in order to be a useful support for different types of analyses: computational and linguistics, as well as for NLP tools training. Arabish, also known as *Arabizi*, is a spontaneous encoding of Arabic dialects in Latin characters and *arithmographs* (numbers used as letters). This **code-system** was developed by Arabic-speaking users of social media in order to facilitate the writing in the Computer Mediated Communication (CMC) and text messaging informal frameworks [[2]](#2). ### *Overview of TArC* In a nutshell, the data gathered in TArC represent **Tunisian Arabish writing and its evolution over the last ten years**. TArC texts have been extracted from social media for an amount of 43 313 tokens. Each text has been extracted together with the user's **metadata** when publicly shared. The metadata consists in: * **The governorate of provenance** * **Age range**: [-25],[25-35],[35-50],[50+] * **Gender**: M/F The Tunisian Arabish texts collected in the TArC have been provided with various annotation levels semi-automatically produced by a Multi-Task Sequence Prediction System: * Token classification into *arabizi*, *foreign* and *emotag*. * Encoding in Arabic Script of the tokens classified as *arabizi* (following the CODA convention [[4]](#4)). * Tokenization of the *CODAfied* texts. * Part-of-Speech tagging for the *arabizi* tokens. TArC numbers: | …