# About Datasets
## L-HSAB Levantine Hate Speech and ABusive
The Levantine Hate Speech and ABusive (L-HSAB) is the first Arabic Levantine Hate Speech and Abusive Language Dataset proposed in the 3rd Workshop ALW-2019 co-located with ACL-2019, Florence, Italy.
In Levantine-speaking countries, especially Syria and Lebanon, the turbulent political/social atmosphere has been linked to intense toxic online debates.
L-HSAB combines 5,846 Syrian/Lebanese political tweets labeled as normal, abusive or hate.
## T-HSAB (Tunisian Hate Speech and ABusive )
The Tunisian Hate Speech and ABusive (T-HSAB) is the first Arabic Tunisian Hate Speech and Abusive Language Dataset proposed in the The 7th International Conference on Arabic Language Processing October 16-17, 2019 (Nancy, France).
T-HSAB combines 6,024 Tunisian comments labeled as normal, abusive or hate.
Hala-Mulki/T-HSAB-A-Tunisia…
## T-SAC
T-SAC is about 17k user comments collected from Facebook users comments written on official pages of Tunisian radios and TV channels namely Mosaique FM, JawhraFM, Shemes FM, HiwarElttounsi TV and Nessma TV.
This data is manually annotated to positive and negative polarities.
For the use of TSAC corpus, please consider the following paper :
hal.archives-ouvertes.fr
## Tunisian Arabizi
On social media, Arabic speakers tend to express themselves in their own local dialect. To do so, Tunisians use ‘Tunisian Arabizi’, where the Latin alphabet is supplemented with numbers.
Tunisian Arabizi combines 70k Tunisian comments scraped by Icompass team (icompass.tn) from different social media platforms and annotated as positive 1 , negative -1 , or neutral 0
## TUNIZI
TUNIZI is collected from comments on social media.
Comments were collected using web scraping techniques to extract comments from YouTube videos.
The dataset is composed of 1500 positive comments and 1500 negative.
Data was collecte …