Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

A Multilingual & Multimodal Text and Image Corpus Dataset for Political Misinformation

Domain:

natural language processingpeace and security

Record type:

dataset
Creator:
DayKadSarAtt
Editor:
Vis
Publisher:
Men
Host:avatar
Our database is a richly annotated multimodal database designed to facilitate strong fake-news detection research. It consists of two complementary but separate components: an image directory and a text spreadsheet. The image directory consists of a folder-level organization with a title as a topic; within each topic directory, the images are then placed in real and fake subdirectories based on the expert labeling. Such an organization allows loading and processing images for cross-modal testing or supervised learning. In contrast, text data are kept in a single Excel sheet where a record is one piece of news. Four separate columns keep the title, source, full news report, and real/fake indicator. Together, these modalities cover a broad range of temporal and topical domains not only social-media posts, mainstream-media news reports, and election-related posts but allowing the training of models on both linguistic aspects (sensational or objective tone, grammaticality, metadata quality) and visual aspects (original vs. photo-manipulated images). With a combination of a sparse folder hierarchy for images and a richly annotated spreadsheet for text, the dataset is well-specified, reproducible, and easy to pipe into any subsequent machine-learning pipeline.

Visit

doi.orgdata.mendeley.com

Tasks

text classification

Tags

PoliticsNatural Language ProcessingMachine LearningMultimodalityConvolutional Neural NetworkPublic SentimentSentiment AnalysisLarge Language ModelMultilingual LLM

Licenses

info:eu-repo/semantics/openAccessCreative Commons Attribution 4.0 Internationalhttps://creativecommons.org/licenses/by/4.0/legalcode

Similar

Emile-Lucky-Muhigira/Multimodal-Image-Text-Misinformation-DetectionAfro SpecDetect A multimodal dataset for African fashion image captioningMultilingual Multimodal Pre-Training with TLI for Zero-Shot Image-Text Retrieval in Low-Resource African LanguagesA Culturally-diverse Multilingual Multimodal Video Benchmark & ModelAfrican Multilingual Text CorpusEgMM-Corpus: A Multimodal Vision-Language Dataset for Egyptian Culture

Emile-Lucky-Muhigira/Multimodal-Image-Text-Misinformation-Detection

Multimodal detection uses image–text consistency to flag misinformation. This study adapts a Fakeddi

Afro SpecDetect A multimodal dataset for African fashion image captioning

Afro SpecDetect A multimodal dataset for African fashion image captioning

Poster presented at the Deep Learning Indaba 2023 by Nouréini Sayouti Souleymane

Multilingual Multimodal Pre-Training with TLI for Zero-Shot Image-Text Retrieval in Low-Resource African Languages

This paper studies zero-shot cross-lingual transfer of vision-language models. Specifically, we focu

A Culturally-diverse Multilingual Multimodal Video Benchmark & Model

Large multimodal models (LMMs) have recently gained attention due to their effectiveness to understa

African Multilingual Text Corpus

NLP model training, translation Notes / challenges: Unequal language representation

EgMM-Corpus: A Multimodal Vision-Language Dataset for Egyptian Culture

Despite recent advances in AI, multimodal culturally diverse datasets are still limited, particularl