Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Babel-ImageNet: Massively Multilingual Evaluation of Vision-and-Language Representations

Domaine:

natural language processing

Type de record:

paperproject

Vision-and-language (VL) models with separate encoders for each modality (e.g., CLIP) have become the go-to models for zero-shot image classification and image-text retrieval. The bulk of the evaluation of these models is, however, performed with English text only: the costly creation of language-specific image-caption datasets has limited multilingual VL benchmarks to a handful of high-resource languages. In this work, we introduce Babel-ImageNet, a massively multilingual benchmark that offers (partial) translations of 1000 ImageNet labels to 92 languages, built without resorting to machine translation (MT) or requiring manual annotation. We instead automatically obtain reliable translations of ImageNext concepts by linking them -- via shared WordNet synsets -- to BabelNet, a massively multilingual lexico-semantic network. We evaluate 8 different publicly available multilingual CLIP models on zero-shot image classification (ZS-IC) for each of the 92 Babel-ImageNet languages, demonstrating a significant gap between English ImageNet performance and that of high-resource languages (e.g., German or Chinese), and an even bigger gap for low-resource languages (e.g., Sinhala or Lao). Crucially, we show that the models' ZS-IC performance on Babel-ImageNet highly correlates with their performance in image-text retrieval, validating that Babel-ImageNet is suitable for estimating the quality of the multilingual VL representation spaces for the vast majority of languages that lack gold image-text data. Finally, we show that the performance of multilingual CLIP for low-resource languages can be drastically improved via cheap, parameter-efficient language-specific training.

Visit

arxiv.orgGithub

Tasks

image-text retrievalimage classificationcomputer vision

Languages

AfrikaansAmharicGaHausaKoMalagasyOromoSomaliSwahiliXhosa

Tags

Vision-and-languagezero-shot image classificationimage-text retrievalBabel-ImageNet

Licenses

BabelNet Non-Commercial License (see https://babelnet.org/full-license)

Similaires

Multilingual Diversity Improves Vision-Language RepresentationsMVL-SIB: A Massively Multilingual Vision-Language Benchmark for Cross-Modal Topical MatchingGlotEval: A Test Suite for Massively Multilingual Evaluation of Large Language ModelsOn the Calibration of Massively Multilingual Language ModelsmSTEB: Massively Multilingual Evaluation of LLMs on Speech and Text TasksSERENGETI: Massively Multilingual Language Models for Africa

Multilingual Diversity Improves Vision-Language Representations

Massive web-crawled image-text datasets lay the foundation for recent progress in multimodal learnin

MVL-SIB: A Massively Multilingual Vision-Language Benchmark for Cross-Modal Topical Matching

Existing multilingual vision-language (VL) benchmarks often only cover a handful of languages. Conse

GlotEval: A Test Suite for Massively Multilingual Evaluation of Large Language Models

Large language models (LLMs) are advancing at an unprecedented pace globally, with regions increasin

On the Calibration of Massively Multilingual Language Models

Massively Multilingual Language Models (MMLMs) have recently gained popularity due to their surprisi

mSTEB: Massively Multilingual Evaluation of LLMs on Speech and Text Tasks

Large Language models (LLMs) have demonstrated impressive performance on a wide range of tasks, incl

SERENGETI: Massively Multilingual Language Models for Africa

Multilingual language models (MLMs) acquire valuable, generalizable linguistic information during pretraining and have advanced the state of the art on task-specific finetuning. So far, only ~ 28 out of ~2,000 African languages are covered in existing language mode