Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Demographical Based Sentiment Analysis for Detection of Hate Speech Tweets for Low Resource Language

Domaine:

natural language processing

Type de record:

dataset
Créateur:
KamShiWasAwa
Éditeur:
Ass
Hôte:
Advancement in IT and communication technology provides the opportunity for social media users to communicate their ideas and thoughts across the globe within no time as well big data promulgated in a result of the communication process itself has immense challenges. Recently, the provision of freedom of speech has witnessed immense promulgation of offensive and hate speech content on the internet aimed the basic human rights violation. The detection of abusive content on social media for rich resource language has become a hot area for researchers in the recent past. However, low-resource languages are underprivileged due to the non-availability of large corpus and its complexity to understand. The proposed methodology mainly has two parts. One is to detect abusive content and the other is to have a demographical analysis of the Indigenously developed dataset. The process starts with the development of a unique unlabeled Urdu dataset of 0.2 M from Twitter through a web scrapper tool named snscraper. The dataset is collected against the 36 districts of Punjab from Pakistan and from the duration 2018- Apr 2022. The dataset is labeled into three target classes Neutral, Offensive, and Hate Speech. After data cleaning, the feature extraction process is achieved with the help of traditional techniques such as Bow and tf-idf with the combination of word and char n-gram and word embedding word2Vec. The dataset is trained on both machine learning algorithms SVM and Logistic regression and deep learning techniques Long Short Term Memory (LSTM) and Convolutional Neural Networks (CNN). The best F score achieved through LSTM on this dataset is 64 and accuracy is 93 through CNN algorithms. A Choropleth map is used for visualization of the dataset distributed among 36 districts of Punjab and a time series plot for time analysis covers five years duration from 2018-Apr to 22.

Visit

doi.org

Tasks

hate speech detectiontext classification

Similaires

Kabila008/Hausa-Tweets-Hate-Speech-Detection-AnalysisDataScience-ArtificialIntelligence/Hate-speech-detection-for-Low-resource-languagesA deep learning based multilingual hate speech detection for resource scarce languagesTweets for sentiment analysisMultilingual Auxiliary Task Scaling for Zero-Shot Hate Speech Detection in Low-Resource LanguagesSOCIAL NETWORK HATE SPEECH DETECTION FOR AMHARIC LANGUAGE

Kabila008/Hausa-Tweets-Hate-Speech-Detection-Analysis

This project involves several complex challenges, from data collection and model development to qual

DataScience-ArtificialIntelligence/Hate-speech-detection-for-Low-resource-languages

# Hate-speech-detection-for-Low-resource-languages ## Overview This project aims to develop machine

A deep learning based multilingual hate speech detection for resource scarce languages

Over the last decade, the increased use of social media has led to an increase in hateful activities

Tweets for sentiment analysis

This is the dataset used in the research manuscript “Sentiment Analysis of Tweets: Political Climate

Multilingual Auxiliary Task Scaling for Zero-Shot Hate Speech Detection in Low-Resource Languages

The goal of hate speech detection is to filter negative online content aiming at certain groups of p

SOCIAL NETWORK HATE SPEECH DETECTION FOR AMHARIC LANGUAGE

The anonymity of social networks makes it attractive for hate speech to mask their criminal activiti