Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Brotherhood at WMT 2024: Leveraging LLM-Generated Contextual Conversations for Cross-Lingual Image Captioning

Domain:

natural language processing

Record type:

paper
Creator:
BetCho
Host:avatar
In this paper, we describe our system under the team name Brotherhood for the English-to-Lowres Multi-Modal Translation Task. We participate in the multi-modal translation tasks for English-Hindi, English-Hausa, English-Bengali, and English-Malayalam language pairs. We present a method leveraging multi-modal Large Language Models (LLMs), specifically GPT-4o and Claude 3.5 Sonnet, to enhance cross-lingual image captioning without traditional training or fine-tuning. Our approach utilizes instruction-tuned prompting to generate rich, contextual conversations about cropped images, using their English captions as additional context. These synthetic conversations are then translated into the target languages. Finally, we employ a weighted prompting strategy, balancing the original English caption with the translated conversation to generate captions in the target language. This method achieved competitive results, scoring 37.90 BLEU on the English-Hindi Challenge Set and ranking first and second for English-Hausa on the Challenge and Evaluation Leaderboards, respectively. We conduct additional experiments on a subset of 250 images, exploring the trade-offs between BLEU scores and semantic similarity across various weighting schemes. Accepted at the Ninth Conference on Machine Translation (WMT24), co-located with EMNLP 2024

Visit

arxiv.org

Tasks

computer visionimage-text retrievalmachine translation

Languages

Hausa

Tags

Computation and LanguageArtificial Intelligence

Similar

gautamiyer31/Image-Captioningamanuelbyte/amharic-image-captioningMuphulusiDzivhani/isiZulu-image-CaptioningNeural Fashion Image Captioning : Accounting for Data DiversityCPE-OOU/NIGERIA-IMAGE-CAPTIONINGkika1s1/afaan-oromo-image-captioning

gautamiyer31/Image-Captioning

A Machine Learning image captioning (image-to-text) Model for three languages – Hausa, Kyrgyz, and

amanuelbyte/amharic-image-captioning

MuphulusiDzivhani/isiZulu-image-Captioning

COS801 Project – isiZulu Image Captioning ## COS 801 Project – Bridging the Visual-Linguistic Divid

Neural Fashion Image Captioning : Accounting for Data Diversity

Image captioning has increasingly large domains of application, and fashion is not an exception. Hav

CPE-OOU/NIGERIA-IMAGE-CAPTIONING

# Nigeria Image Captioning This repository contains the code and resources for our final year proje

kika1s1/afaan-oromo-image-captioning

Generate meaningful captions in Afaan Oromo from images using CNN encoders and Transformer decoders