Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Align Where the Words Look: Cross-Attention-Guided Patch Alignment with Contrastive and Transport Regularization for Bengali Captioning

Domaine:

natural language processing

Type de record:

paperdatasetmodel
Créateur:
AnoZabRah
Hôte:avatar
Grounding vision--language models in low-resource languages remains challenging, as they often produce fluent text about the wrong objects. This stems from scarce paired data, translation pivots that break alignment, and English-centric pretraining that ignores target-language semantics. We address this with a compute-aware Bengali captioning pipeline trained on LaBSE-verified EN--BN pairs and 110k bilingual-prompted synthetic images. A frozen MaxViT yields stable visual patches, a Bengali-native mBART-50 decodes, and a lightweight bridge links the modalities. Our core novelty is a tri-loss objective: Patch-Alignment Loss (PAL) aligns real and synthetic patch descriptors using decoder cross-attention, InfoNCE enforces global real--synthetic separation, and Sinkhorn-based OT ensures balanced fine-grained patch correspondence. This PAL+InfoNCE+OT synergy improves grounding, reduces spurious matches, and drives strong gains on Flickr30k-1k (BLEU-4 12.29, METEOR 27.98, BERTScore-F1 71.20) and MSCOCO-1k (BLEU-4 12.00, METEOR 28.14, BERTScore-F1 75.40), outperforming strong CE baselines and narrowing the real--synthetic centroid gap by 41%.

Visit

arxiv.org

Tasks

image-text retrievalcomputer vision

Tags

Computer Vision and Pattern RecognitionArtificial Intelligence

Similaires

Align to the Pivot: Dual Alignment with Self-Feedback for Multilingual Math ReasoningContrastive Learning for Cross-Lingual Alignment and Robustness in Multimodal ModelsToken-Region Guided Cross-Attention Fusion for Multimodal Affect InterpretationBengali Image Captioning Using Vision Encoder-Decoder ModelOptimal Transport Alignment for Reducing Cross-Lingual Retrieval Performance GapsOptimal Transport Distillation for Cross-Lingual Performance Alignment in XQuAD

Align to the Pivot: Dual Alignment with Self-Feedback for Multilingual Math Reasoning

Despite the impressive reasoning abilities demonstrated by large language models (LLMs), empirical e

Contrastive Learning for Cross-Lingual Alignment and Robustness in Multimodal Models

Pre-trained multilingual language encoders, such as multilingual BERT and XLM-R, show great potentia

Token-Region Guided Cross-Attention Fusion for Multimodal Affect Interpretation

Automated analysis of multimodal content on social networks has become a critical task for understan

Bengali Image Captioning Using Vision Encoder-Decoder Model

Optimal Transport Alignment for Reducing Cross-Lingual Retrieval Performance Gaps

Benefiting from transformer-based pre-trained language models, neural ranking models have made signi

Optimal Transport Distillation for Cross-Lingual Performance Alignment in XQuAD

Benefiting from transformer-based pre-trained language models, neural ranking models have made signi