Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Comparative Performance of Ensemble Machine Learning for Arabic Cyberbullying and Offensive Language Detection

Domain:

natural language processing

Record type:

paper
Creator:
MarTarAhmTar
Publisher:
Spr
Host:
Abstract In recent years, research on abusive language and cyberbullying detection have gained a great deal of interest since it affects both individual victims and societies. Hateful communications, bullying, sexism, racism, aggressive content, harassment, toxic comments, and other forms of abuse have all increased dramatically due to the ease of access to social media platforms such as Facebook, Instagram, Twitter, and others. As a result, there is a significant need to identify, manage, and restrict the spread of offensive content on social networking sites, prompting us to perform this study to automate the detection of offensive language or cyberbullying. Having a balanced data set for a model would generate higher accuracy models, thus we build a new Arabic balanced data set to be used in the process of offensive detection. Lately, Ensemble Machine Learning has been used to enhance the performance of single classifiers. The aim of this study is to compare the performance of different single and ensemble machine learning algorithms in detecting Arabic text containing cyberbullying and offensive language. For this purpose, we have chosen three machine learning classifiers and three ensemble models and apply them to three Arabic datasets two of them are offensive datasets that are publicly available, and the third one which we constructed. The results showed that the ensemble machine learning methodology outperforms the single learner machine learning approach. Voting performs is the best among the trained ensemble machine learning classifiers, with accuracy scores of (71.1%, 76.7%, and 98.5%) for the three used datasets respectively, exceeding the score obtained by the best single learner classifier (65.1%, 76.2%, and 98%) for the same datasets. Finally we use hyperparameters tunning on the Arabic cyberbylluing data set to optize the performance of the voting technique.

Visit

doi.org

Tasks

hate speech detectiontext classification

Licenses

https://creativecommons.org/licenses/by/4.0/