Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Improving the Detection of Multilingual South African Abusive Language via Skip-gram Using Joint Multilevel Domain Adaptation

Domain:

natural language processing

Record type:

paper
Creator:
OluEdu
Publisher:
Ass
Host:
The distinctiveness and sparsity of low-resource multilingual South African abusive language necessitate the development of a novel solution to automatically detect different classes of abusive language instances using machine learning. Skip-gram has been used to address sparsity in machine learning classification problems but is inadequate in detecting South African abusive language due to the considerable amount of rare features and class imbalance. Joint Domain Adaptation has been used to enlarge features of a low-resource target domain for improved classification outcomes by jointly learning from the target domain and large-resource source domain. This article, therefore, builds a Skip-gram model based on Joint Domain Adaptation to improve the detection of multilingual South African abusive language. Contrary to the existing Joint Domain Adaptation approaches, a Joint Multilevel Domain Adaptation model involving adaptation of monolingual source domain instances and multilingual target domain instances with high frequency of rare features was executed at the first level and adaptation of target-domain features and first-level features at the next level. Both surface-level and embedding word features were used to evaluate the proposed model. In the evaluation of surface-level features, the Joint Multilevel Domain Adaptation model outperformed the state-of-the-art models with accuracy of 0.92 and F1-score of 0.68. In the evaluation of embedding features, the proposed model outperformed the state-of-the-art models with accuracy of 0.88 and F1-score of 0.64. The Joint Multilevel Domain Adaptation model significantly improved the average information gain of the rare features in different language categories and reduced class imbalance.

Visit

doi.org

Tasks

hate speech detectiontext classification

Licenses

https://www.acm.org/publications/policies/copyright_policy#Background

Similar

Efficient multilingual and domain adaptation of language models under resource constraintsImproved semi-supervised learning technique for automatic detection of South African abusive language on TwitterTigrinya Abusive Language Detection (TiALD) DatasetPysham0n/Tunisian-Dialect-Abusive-Language-DetectionAfriHate: A Multilingual Collection of Hate Speech and Abusive Language Datasets for African LanguagesMassively Multilingual Adaptation of Large Language Models Using Bilingual Translation Data

Efficient multilingual and domain adaptation of language models under resource constraints

Neural networks trained for language modeling, which is the task of finding the missing words in a g

Improved semi-supervised learning technique for automatic detection of South African abusive language on Twitter

Semi-supervised learning is a potential solution for improving training data in low-resourced abusiv

Tigrinya Abusive Language Detection (TiALD) Dataset

TiALD is a large-scale, multi-task benchmark dataset for abusive language detection in the Tigrinya

Pysham0n/Tunisian-Dialect-Abusive-Language-Detection

Abusive language detector in comments written in the Tunisian dialect. Using a combination of web sc

AfriHate: A Multilingual Collection of Hate Speech and Abusive Language Datasets for African Languages

Hate speech and abusive language are global phenomena that need socio-cultural background knowledge

Massively Multilingual Adaptation of Large Language Models Using Bilingual Translation Data

This paper investigates a critical design decision in the practice of massively multilingual continu