Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Brotherhood at WMT 2024: Leveraging LLM-Generated Contextual Conversations for Cross-Lingual Image Captioning

Domaine:

natural language processing

Type de record:

paper
Créateur:
BetCho
Hôte:avatar
In this paper, we describe our system under the team name Brotherhood for the English-to-Lowres Multi-Modal Translation Task. We participate in the multi-modal translation tasks for English-Hindi, English-Hausa, English-Bengali, and English-Malayalam language pairs. We present a method leveraging multi-modal Large Language Models (LLMs), specifically GPT-4o and Claude 3.5 Sonnet, to enhance cross-lingual image captioning without traditional training or fine-tuning. Our approach utilizes instruction-tuned prompting to generate rich, contextual conversations about cropped images, using their English captions as additional context. These synthetic conversations are then translated into the target languages. Finally, we employ a weighted prompting strategy, balancing the original English caption with the translated conversation to generate captions in the target language. This method achieved competitive results, scoring 37.90 BLEU on the English-Hindi Challenge Set and ranking first and second for English-Hausa on the Challenge and Evaluation Leaderboards, respectively. We conduct additional experiments on a subset of 250 images, exploring the trade-offs between BLEU scores and semantic similarity across various weighting schemes. Accepted at the Ninth Conference on Machine Translation (WMT24), co-located with EMNLP 2024

Visit

arxiv.org

Tasks

computer visionimage-text retrievalmachine translation

Languages

Hausa

Tags

Computation and LanguageArtificial Intelligence

Similaires

gautamiyer31/Image-Captioningamanuelbyte/amharic-image-captioningMuphulusiDzivhani/isiZulu-image-CaptioningNeural Fashion Image Captioning : Accounting for Data DiversityCPE-OOU/NIGERIA-IMAGE-CAPTIONINGkika1s1/afaan-oromo-image-captioning

gautamiyer31/Image-Captioning

A Machine Learning image captioning (image-to-text) Model for three languages – Hausa, Kyrgyz, and

amanuelbyte/amharic-image-captioning

MuphulusiDzivhani/isiZulu-image-Captioning

COS801 Project – isiZulu Image Captioning ## COS 801 Project – Bridging the Visual-Linguistic Divid

Neural Fashion Image Captioning : Accounting for Data Diversity

Image captioning has increasingly large domains of application, and fashion is not an exception. Hav

CPE-OOU/NIGERIA-IMAGE-CAPTIONING

# Nigeria Image Captioning This repository contains the code and resources for our final year proje

kika1s1/afaan-oromo-image-captioning

Generate meaningful captions in Afaan Oromo from images using CNN encoders and Transformer decoders