Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Retrieval Augmented Enhanced Dual Co-Attention Framework for Target Aware Multimodal Bengali Hateful Meme Detection

Domaine:

natural language processing

Type de record:

paperdatasetmodel
Créateur:
TanAla
Hôte:avatar
Hateful content on social media increasingly appears as multimodal memes that combine images and text to convey harmful narratives. In low-resource languages such as Bengali, automated detection remains challenging due to limited annotated data, class imbalance, and pervasive code-mixing. To address these issues, we augment the Bengali Hateful Memes (BHM) dataset with semantically aligned samples from the Multimodal Aggression Dataset in Bengali (MIMOSA), improving both class balance and semantic diversity. We propose the Enhanced Dual Co-attention Framework (xDORA), integrating vision encoders (CLIP, DINOv2) and multilingual text encoders (XGLM, XLM-R) via weighted attention pooling to learn robust cross-modal representations. Building on these embeddings, we develop a FAISS-based k-nearest neighbor classifier for non-parametric inference and introduce RAG-Fused DORA, which incorporates retrieval-driven contextual reasoning. We further evaluate LLaVA under zero-shot, few-shot, and retrieval-augmented prompting settings. Experiments on the extended dataset show that xDORA (CLIP + XLM-R) achieves macro-average F1-scores of 0.78 for hateful meme identification and 0.71 for target entity detection, while RAG-Fused DORA improves performance to 0.79 and 0.74, yielding gains over the DORA baseline. The FAISS-based classifier performs competitively and demonstrates robustness for rare classes through semantic similarity modeling. In contrast, LLaVA exhibits limited effectiveness in few-shot settings, with only modest improvements under retrieval augmentation, highlighting constraints of pretrained vision-language models for code-mixed Bengali content without fine-tuning. These findings demonstrate the effectiveness of supervised, retrieval-augmented, and non-parametric multimodal frameworks for addressing linguistic and cultural complexities in low-resource hate speech detection.

Visit

arxiv.org

Tasks

hate speech detectiontext classification

Tags

Computation and Language

Similaires

Error-Aware TF-IDF Retrieval-Augmented Generation for ASR Error CorrectionAlign before Attend: Aligning Visual and Textual Features for Multimodal Hateful Content DetectionMulti-stream Attention-Enhanced Deep Learning Framework for Cocoa Leaf Disease Detection and Classification in GhanaA Context-Aware and Target-Adaptive Multilingual Framework for Hate Speech Detection in Code-Switched Social Media TextRATIO: A Graph-Aware Framework for Precedent-Grounded Legal RetrievalAn Evidence-Grounded Retrieval-Augmented Transformer Framework for Health Misinformation Verification

Error-Aware TF-IDF Retrieval-Augmented Generation for ASR Error Correction

End-to-end automatic speech recognition systems frequently hallucinate rare entities and domain-spec

Align before Attend: Aligning Visual and Textual Features for Multimodal Hateful Content Detection

Multimodal hateful content detection is a challenging task that requires complex reasoning across vi

Multi-stream Attention-Enhanced Deep Learning Framework for Cocoa Leaf Disease Detection and Classification in Ghana

A Context-Aware and Target-Adaptive Multilingual Framework for Hate Speech Detection in Code-Switched Social Media Text

The rapid expansion of social media has accelerated the spread of hate speech, particularly within m

RATIO: A Graph-Aware Framework for Precedent-Grounded Legal Retrieval

A graph-aware retrieval framework for legal precedent search that explicitly models the citation net

An Evidence-Grounded Retrieval-Augmented Transformer Framework for Health Misinformation Verification

The rapid spread of false and misleading health information through digital platforms has become a m