Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Improved semi-supervised learning technique for automatic detection of South African abusive language on Twitter

Domaine:

natural language processing

Type de record:

paper
Créateur:
OluEdu
Éditeur:
Sou
Hôte:
Semi-supervised learning is a potential solution for improving training data in low-resourced abusive language detection contexts such as South African abusive language detection on Twitter. However, the existing semi-supervised learning methods have been skewed towards small amounts of labelled data, with small feature space. This paper, therefore, presents a semi-supervised learning technique that improves the distribution of training data by assigning labels to unlabelled data based on the majority voting over different feature sets of labelled and unlabelled data clusters. The technique is applied to South African English corpora consisting of labelled and unlabelled abusive tweets. The proposed technique is compared with state-of-the-art self-learning and active learning techniques based on syntactic and semantic features. The performance of these techniques with Logistic Regression, Support Vector Machine and Neural Networks are evaluated. The proposed technique, with accuracy and F1-score of 0.97 and 0.95, respectively, outperforms existing semi-supervised learning techniques. The learning curves show that the training data was used more efficiently by the proposed technique compared to existing techniques. Overall, n-gram syntactic features with a Logistic Regression classifier records the highest performance. The paper concludes that the proposed semi-supervised learning technique effectively detected implicit and explicit South African abusive language on Twitter.

Visit

doi.org

Tasks

hate speech detectiontext classification

Licenses

http://creativecommons.org/licenses/by-nc/4.0

Similaires

Image Segmentation for Dust Detection Using Semi-supervised Machine LearningSemi-supervised learning approaches for predicting South African political sentiment for local government electionsA Supervised Learning Model for the Automatic Assessment of Language Levels Based on Learner ErrorsSemi-Supervised Anomaly Detection for the Determination of Vehicle Hijacking TweetsTigrinya Abusive Language Detection (TiALD) DatasetPysham0n/Tunisian-Dialect-Abusive-Language-Detection

Image Segmentation for Dust Detection Using Semi-supervised Machine Learning

Dust plumes originating from the Earth’s major arid and semi-arid areas can significantly affect the

Semi-supervised learning approaches for predicting South African political sentiment for local government elections

This study aims to understand the South African political context by analysing the sentiments shared

A Supervised Learning Model for the Automatic Assessment of Language Levels Based on Learner Errors

International audience This paper focuses on the use of technology in language learni

Semi-Supervised Anomaly Detection for the Determination of Vehicle Hijacking Tweets

In South Africa, there is an ever-growing issue of vehicle hijackings. This leads to travellers cons

Tigrinya Abusive Language Detection (TiALD) Dataset

TiALD is a large-scale, multi-task benchmark dataset for abusive language detection in the Tigrinya

Pysham0n/Tunisian-Dialect-Abusive-Language-Detection

Abusive language detector in comments written in the Tunisian dialect. Using a combination of web sc