Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Token-Region Guided Cross-Attention Fusion for Multimodal Affect Interpretation

Domaine:

natural language processing

Type de record:

paper
Créateur:
FarRez
Éditeur:
arXiv
Hôte:avatar
Automated analysis of multimodal content on social networks has become a critical task for understanding public sentiment and information diffusion in the digital age. However, classifying internet memes remains computationally challenging due to the intricate interplay between visual cues and embedded, often stylized, text, particularly in low-resource languages like Bengali Language. This paper addresses the detection of political intent in Bengali memes by introducing Multimodal Cross-Attention Fusion framework. We first leverage a Vision-Language Model to extract high-fidelity OCR text from noisy meme images. Subsequently, we encode visual and textual features and synthesize them through a cross-modal multi-head attention mechanism that aligns semantic tokens with visual regions. We also investigate the integration of a domain-specific political lexicon as a knowledge prior. Experimental evaluation on the PoliMemeDecode1 dataset shows that our attention-based fusion significantly outperforms unimodal baselines and standard concatenation methods, achieving a state-of-the-art Macro-F1 of approximately 0.94. Interpretability analyzes further confirm that the model effectively learns to ground textual semantics in visual evidence.

Visit

doi.org

Tags

Computer Vision and Pattern Recognition (cs.CV)Artificial Intelligence (cs.AI)FOS: Computer and information sciences

Licenses

Creative Commons Attribution 4.0 Internationalhttps://creativecommons.org/licenses/by/4.0/legalcode

Similaires

MulMoSenT: Multimodal Sentiment Analysis for a Low-Resource Language Using Textual-Visual Cross-Attention and FusionCLIF-Net: Intersection-guided Cross-view Fusion Network for Infection Detection from Cranial UltrasoundMultimodal Fusion Strategies for Mapping Biophysical Landscape FeaturesA Multimodal AI Copilot for Obstetric Ultrasound InterpretationFusion Complexity Inversion: Why Simpler Cross View Modules Outperform SSMs and Cross View Attention Transformers for Pasture Biomass RegressionHassan48khan/SavorNet-Adaptive-Attention-Fusion-for-Ethiopian-Cuisine-Classification

MulMoSenT: Multimodal Sentiment Analysis for a Low-Resource Language Using Textual-Visual Cross-Attention and Fusion

First-ever Bengali Multimodal Sentiment Analysis (BMSA) corpus and details in https://www.sciencedir

CLIF-Net: Intersection-guided Cross-view Fusion Network for Infection Detection from Cranial Ultrasound

Abstract This paper addresses the problem of detecting possible serious bacterial

Multimodal Fusion Strategies for Mapping Biophysical Landscape Features

Multimodal aerial data are used to monitor natural systems, and machine learning can significantly a

A Multimodal AI Copilot for Obstetric Ultrasound Interpretation

Ultrasound interpretation in Rwanda is largely confined to tertiary facilities: approximately 115 ob

Fusion Complexity Inversion: Why Simpler Cross View Modules Outperform SSMs and Cross View Attention Transformers for Pasture Biomass Regression

Accurate estimation of pasture biomass from agricultural imagery is critical for sustainable livesto

Hassan48khan/SavorNet-Adaptive-Attention-Fusion-for-Ethiopian-Cuisine-Classification

SavorNet: Adaptive Attention Fusion for Ethiopian Cuisine Classification Overview SavorNet is a deep