Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

BOUTEF: Bolstering Our Understanding Through an Elaborated Fake News Corpus

Domaine:

natural language processing

Type de record:

paperdataset
Créateur:
SmaHamDavDje
Éditeur:
StaAnaAgeANR
Éditeur:
CCSD
Hôte:avatar
International audience This article presents BOUTEF, an original and comprehensive corpus of fake news. It encompassescontent in Algerian and Tunisian dialects, Modern Standard Arabic (MSA), French, and English,featuring instances of code-switching between these languages. Moreover, for the Algerian and Tunisiandialects, we have preserved both Latin and Arabic scripts in the dataset. BOUTEF comprises over 3,600fake news posts collected from various social media platforms spanning from 2010 to 2024. This corpusis developed as part of the TRADEF 4 project and is made available to the research community. Eachfake news post in BOUTEF is associated with 16 attributes, providing rich contextual information. Thedata was gathered from Facebook, Twitter, YouTube, and TikTok, reflecting the diverse sources of misinformation.To enhance the depth of our analysis, we introduce a novel labeling scheme consisting of 40categories. This scheme is developed through a thorough examination of the collected corpus, and we havealso retained a tagging process inspired by Claire Wardle’s categorization. BOUTEF not only contributesto the understanding of fake news in multilingual contexts but also provides valuable resources for furtherresearch in this domain.

Visit

hal.science

Tasks

code switchingtext classification

Languages

Arabic, Algerian SpokenArabic, Tunisian Spoken

Tags

Fake news Multilingual Context Arabic dialect Code-switching Social Media misinformation Labeling categoriesFake newsMultilingual ContextArabic dialectCode-switchingSocial Media misinformationLabeling categories[INFO]Computer Science [cs]

Licenses

https://about.hal.science/hal-authorisation-v1/info:eu-repo/semantics/OpenAccess

Similaires

BOUTEF: A Multilingual Corpus for Fake News in North Africa - Language as a WeaponBLUFF: Benchmark for Linguistic Understanding of Fake-news ForensicsThe First Corpus for Detecting Fake News in Hausa LanguageFASSILA: A Corpus for Algerian Dialect Fake News Detection and Sentiment AnalysisUnderstanding the Impact of and Analysing Fake News About COVID-19 in SAkalkidanyishak/fake-news-data

BOUTEF: A Multilingual Corpus for Fake News in North Africa - Language as a Weapon

The rapid spread of fake news on social media has become a major challenge, p

BLUFF: Benchmark for Linguistic Understanding of Fake-news Forensics

BLUFF is a comprehensive multilingual benchmark for fake news detection spanning 79 languages with o

The First Corpus for Detecting Fake News in Hausa Language

FASSILA: A Corpus for Algerian Dialect Fake News Detection and Sentiment Analysis

In the context of low-resource languages, the Algerian dialect (AD) faces challenges due to the abse

Understanding the Impact of and Analysing Fake News About COVID-19 in SA

kalkidanyishak/fake-news-data

fake news data aggregation for Amharic # የአማርኛ NLP የመረጃ ስብስብ ቁልፍ ግንዛቤዎች **ለመረጃዎ ግንዛቤን ለማግኘት እና ለሞዴ