Named Entity Recognition for Algerian
# NERDz -- Named Entity Recognition for Algerian
This dataset is described in the paper "*NERDz: A Preliminary Dataset of Named Entities for Algerian*" by Samia Touileb, published in the Proceedings of the 2nd Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics and the 12th International Joint Conference on Natural Language Processing (Volume 2: Short Papers).
In the NERDz dataset we added named entity annotations on top of the extension of the NArabizi treebank presented by *Touileb and Barnes (2021)*, and which was originally introduced by *Seddah et al. (2020)*.
# The NArabizi treebank
The NArabizi treebank was developed by Seddah et al. (2020). It contains manually annotated syntactic and morphological information, and comprises around 1,500 sentences written in the Algerian dialect. These are mostly comments from newspapers' web forums (1,300 sentences from (Cotterell et al., 2014)), in addition to 200 sentences from song lyrics. The sentences are annotated on five different levels, covering tokenization, morphology, identification of code-switching, syntax, and translation to French (Seddah et al., 2020). Touileb and Barnes (2021) have further extended the NArabizi treebank, by first cleaning the treebank for duplicates, correcting some of the French translations, and some of the code-switching labels. But most importantly, they manually transliterated each sentence into purely Arabic script and code-switched scripts. The treebank therefore has three parallel writing forms for each token in a sentence. They have also annotated each sentence of the treebank for sentiment and topic (see Touileb and Barnes (2021) for more details). Due to the preprocessing, this version of the treebank (Touileb and Barnes, 2021) is a little bit smaller than the original treebank (Seddah et al., 2020). NERDz is built on top of this modified version.
# The Named Entity annotations
The named entity annotations in NERDz are continuous, no …