Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

A Multilingual & Multimodal Text and Image Corpus Dataset for Political Misinformation

Domaine:

natural language processingpeace and security

Type de record:

dataset
Créateur:
DayKadSarAtt
Éditeur:
Vis
Éditeur:
Men
Hôte:avatar
Our database is a richly annotated multimodal database designed to facilitate strong fake-news detection research. It consists of two complementary but separate components: an image directory and a text spreadsheet. The image directory consists of a folder-level organization with a title as a topic; within each topic directory, the images are then placed in real and fake subdirectories based on the expert labeling. Such an organization allows loading and processing images for cross-modal testing or supervised learning. In contrast, text data are kept in a single Excel sheet where a record is one piece of news. Four separate columns keep the title, source, full news report, and real/fake indicator. Together, these modalities cover a broad range of temporal and topical domains not only social-media posts, mainstream-media news reports, and election-related posts but allowing the training of models on both linguistic aspects (sensational or objective tone, grammaticality, metadata quality) and visual aspects (original vs. photo-manipulated images). With a combination of a sparse folder hierarchy for images and a richly annotated spreadsheet for text, the dataset is well-specified, reproducible, and easy to pipe into any subsequent machine-learning pipeline.

Visit

doi.orgdata.mendeley.com

Tasks

text classification

Tags

PoliticsNatural Language ProcessingMachine LearningMultimodalityConvolutional Neural NetworkPublic SentimentSentiment AnalysisLarge Language ModelMultilingual LLM

Licenses

info:eu-repo/semantics/openAccessCreative Commons Attribution 4.0 Internationalhttps://creativecommons.org/licenses/by/4.0/legalcode

Similaires

Emile-Lucky-Muhigira/Multimodal-Image-Text-Misinformation-DetectionAfro SpecDetect A multimodal dataset for African fashion image captioningMultilingual Multimodal Pre-Training with TLI for Zero-Shot Image-Text Retrieval in Low-Resource African LanguagesA Culturally-diverse Multilingual Multimodal Video Benchmark & ModelAfrican Multilingual Text CorpusEgMM-Corpus: A Multimodal Vision-Language Dataset for Egyptian Culture

Emile-Lucky-Muhigira/Multimodal-Image-Text-Misinformation-Detection

Multimodal detection uses image–text consistency to flag misinformation. This study adapts a Fakeddi

Afro SpecDetect A multimodal dataset for African fashion image captioning

Afro SpecDetect A multimodal dataset for African fashion image captioning

Poster presented at the Deep Learning Indaba 2023 by Nouréini Sayouti Souleymane

Multilingual Multimodal Pre-Training with TLI for Zero-Shot Image-Text Retrieval in Low-Resource African Languages

This paper studies zero-shot cross-lingual transfer of vision-language models. Specifically, we focu

A Culturally-diverse Multilingual Multimodal Video Benchmark & Model

Large multimodal models (LMMs) have recently gained attention due to their effectiveness to understa

African Multilingual Text Corpus

NLP model training, translation Notes / challenges: Unequal language representation

EgMM-Corpus: A Multimodal Vision-Language Dataset for Egyptian Culture

Despite recent advances in AI, multimodal culturally diverse datasets are still limited, particularl