Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Misinformation detection in Luganda-English code-mixed social media text

Domain:

natural language processing

Record type:

datasetpaper
Creator:
NabKabBabirye, ClaireTus
Publisher:
arXiv
Host:avatar
The increasing occurrence, forms, and negative effects of misinformation on social media platforms has necessitated more misinformation detection tools. Currently, work is being done addressing COVID-19 misinformation however, there are no misinformation detection tools for any of the 40 distinct indigenous Ugandan languages. This paper addresses this gap by presenting basic language resources and a misinformation detection data set based on code-mixed Luganda-English messages sourced from the Facebook and Twitter social media platforms. Several machine learning methods are applied on the misinformation detection data set to develop classification models for detecting whether a code-mixed Luganda-English message contains misinformation or not. A 10-fold cross validation evaluation of the classification methods in an experimental misinformation detection task shows that a Discriminative Multinomial Naive Bayes (DMNB) method achieves the highest accuracy and F-measure of 78.19% and 77.90% respectively. Also, Support Vector Machine and Bagging ensemble classification models achieve comparable results. These results are promising since the machine learning models are based on n-gram features from only the misinformation detection dataset. Accepted at African NLP workshop @EACL 2021

Visit

doi.orgarxiv.org

Tasks

code switchingtext classification

Languages

Ganda

Tags

Computation and Language (cs.CL)Social and Information Networks (cs.SI)FOS: Computer and information sciencesFOS: Computer and information sciences

Licenses

Creative Commons Attribution 4.0 Internationalhttps://creativecommons.org/licenses/by/4.0/legalcode

Similar

Abusive Content Detection in Telugu-English Code-Mixed Social Media Using Hybrid Transformer ArchitecturesA Comparative Study of Transformer-based Models for Hate-Speech Detection in English-Kiswahili Code-Switched Social Media TextMultimodal Misinformation Detection in a South African Social Media Environmentlichmond/MULTILINGUAL-MISINFORMATION-DETECTION-IN-SOCIAL-MEDIA-AND-THE-INTERNETDetecting Propaganda Techniques in Code-Switched Social Media TextFine-Tuning Large Language Models with QLoRA for Offensive Language Detection in Roman Urdu-English Code-Mixed Text

Abusive Content Detection in Telugu-English Code-Mixed Social Media Using Hybrid Transformer Architectures

The rapid growth of social media platforms has led to a substantial increase in user-generated conte

A Comparative Study of Transformer-based Models for Hate-Speech Detection in English-Kiswahili Code-Switched Social Media Text

The transformer architecture, first introduced in 2017 by researchers at Google, has revolutionized

Multimodal Misinformation Detection in a South African Social Media Environment

With the constant spread of misinformation on social media networks, a need has arisen to continuous

lichmond/MULTILINGUAL-MISINFORMATION-DETECTION-IN-SOCIAL-MEDIA-AND-THE-INTERNET

Repo for the research paper "Multilingual Misinformation Detection In Social Media And The Internet"

Detecting Propaganda Techniques in Code-Switched Social Media Text

Propaganda is a form of communication intended to influence the opinions and the mindset of the publ

Fine-Tuning Large Language Models with QLoRA for Offensive Language Detection in Roman Urdu-English Code-Mixed Text

The use of derogatory terms in languages that employ code mixing, such as Roman Urdu, presents chall