Comprehensive list of resources for automated processing of Tunisian dialect text.
## Introduction
The Tunisian dialect is an under-resourced language. This project is a personal effort to put scattered sources of information,
resources, and tools related to automated processing of text written in the Tunisian dialect in one place for all to use.
Hopefully it will save NLP practitioners valuable time spent otherwise chasing after information scattered all over the Web.
I have categorized this information into the following categories:
1. Publicly Available Corpora and Lexicons
2. NLP Software
3. Scientific Papers and Articles
4. Books
5. Web Articles & Links
6. Academic research groups & labs
7. Conferences & Workshops
If you would like to contribute or add a new listing, please either send a git pull request or email me at chiraz.benabdelkader@gmail.com
## Publicly Available Corpora and Lexicons
- The MADAR Arabic Dialect Corpus and Lexicon, H. Bouamor et al.,
nlp.qatar.cmu.edu
Description:
"The latest version of the lexicon is available for browsing online."
- Tunisian Arabic Corpus, by Karen McNeil and Miled Faiza,
tunisiya.org
Description:
"There are currently 2,006 texts in the corpus, comprising 881,964 words.
The main categories currently included are
1) traditional written sources (folklore, songs, folk poetry, proverb collections, screenplays)
2) new written sources (blogs, email, Facebook, forum postings) -- currently the dominant source
3) transcribed audio (e.g. from radio programming)."
"The corpus is available freely online at tunisiya.org, where users can perform complex concordance
searches and view search results in context, with access to the full text. "
- DID-LREC-2018: Training and test data for the Arabic dialect identification (DID) shared task at LREC 2018,
github.com
- Various small lexicons contributed by N. Karmani Ben Moussa as part of her PhD thesis work, last updated August 2016,
github.com
- AOC …