Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Kaberia-Timo/Multi-Task-Classification-of-Kiswahili-Online-Misinformation-and-Harmful-language

Domaine:

natural language processing

Type de record:

projectsoftware
Créateur:
Kab
Hôte:
\# Multi-task Misinformation Detection for Low-Resource Swahili A reproducible deep learning pipeline for multi-task learning on \*\*Swahili misinformation detection\*\* and \*\*Swahili hate speech detection\*\* using shared transformer representations. This repository accompanies the research project investigating whether \*\*joint multi-task learning\*\* can improve misinformation detection by leveraging knowledge from a related hate speech detection task in a low-resource language. \--- \## Overview The project develops and evaluates two multi-task transformer models: 1\. \*\*Baseline Multi-task Model\*\* - Equal task weighting - Joint learning of misinformation and hate speech 2\. \*\*Misinformation-Tuned Multi-task Model\*\* - Increased optimisation emphasis on misinformation detection - Same architecture as the baseline - Different task-loss weighting strategy The experiments use: \- \*\*PolitikWeli\*\* (Swahili misinformation dataset) \- \*\*AfriHate\*\* (Swahili hate speech dataset) Both datasets are harmonised into a unified multi-task training dataset while preventing data leakage and ensuring reproducibility. \--- \## Repository Structure ```text multitask-misinformation-detection/ │ ├── data/ │ ├── raw/ │ └── processed/ │ ├── notebooks/ │ ├── 01\_politikweli\_prepare.ipynb │ ├── 02\_afrihate\_prepare.ipynb │ ├── 03\_build\_multitask\_dataset.ipynb │ ├── 04\_baseline\_training.ipynb │ └── 05\_misinformation\_tuned\_training.ipynb │ ├── models/ │ ├── baseline/ │ └── misinformation\_tuned/ │ ├── reports/ ├── docs/ └── src/ ``` \--- \## Workflow The notebooks should be executed in the following order. \### Notebook 1 \*\*01\_politikweli\_prepare.ipynb\*\* Purpose \- Load hydrated PolitikWeli datasets \- Clean and harmonise records \- Construct binary misinformation labels \- Perform stratified train/validation/te …

Visit

github.com

Tasks

hate speech detectiontext classification

Languages

SwahiliSwahili, CoastalSwahili, Congo

Licenses

MIT

Similaires

A multi-task learning framework for sentiment analysis and news classification for low-resource languageTask Mate Kenyan Sign Language Classification ChallengeResources Building for Arabic Harmful Online Content: SurveyEfficient Multi-Task Arabic Question Classification with a Shared MARBERT EncoderCOVID-19-related online misinformation in BangladeshHitachi at SemEval-2023 Task 3: Exploring Cross-lingual Multi-task Strategies for Genre and Framing Detection in Online News

A multi-task learning framework for sentiment analysis and news classification for low-resource language

Despite the growing progress in Natural Language Processing (NLP), low-resource languages such as Ha

Task Mate Kenyan Sign Language Classification Challenge

Can you classify words in Kenyan Sign Language?
The data was collected by 800 taskers from Kenya, Mexico and India. There are nine classes, each a different sign.
The objective of this competition is to classify the ten different Sign Language signs pres

Resources Building for Arabic Harmful Online Content: Survey

International audience

Users of social networks and Internet sites face numer

Efficient Multi-Task Arabic Question Classification with a Shared MARBERT Encoder

No description provided. If you use this software, please cite the associated paper.

COVID-19-related online misinformation in Bangladesh

Purpose This paper aims to understand the popular themes of coronavirus disease 2019 (COVID-19)-rel

Hitachi at SemEval-2023 Task 3: Exploring Cross-lingual Multi-task Strategies for Genre and Framing Detection in Online News

This paper explains the participation of team Hitachi to SemEval-2023 Task 3 "Detecting the genre, t