Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

CAPIVARA: Cost-Efficient Approach for Improving Multilingual CLIP Performance on Low-Resource Languages

Domaine:

natural language processing

Type de record:

papermodel
Créateur:
SanMorFerSil
Hôte:avatar
This work introduces CAPIVARA, a cost-efficient framework designed to enhance the performance of multilingual CLIP models in low-resource languages. While CLIP has excelled in zero-shot vision-language tasks, the resource-intensive nature of model training remains challenging. Many datasets lack linguistic diversity, featuring solely English descriptions for images. CAPIVARA addresses this by augmenting text data using image captioning and machine translation to generate multiple synthetic captions in low-resource languages. We optimize the training pipeline with LiT, LoRA, and gradient checkpointing to alleviate the computational cost. Through extensive experiments, CAPIVARA emerges as state of the art in zero-shot tasks involving images and Portuguese texts. We show the potential for significant improvements in other low-resource languages, achieved by fine-tuning the pre-trained multilingual CLIP using CAPIVARA on a single GPU for 2 hours. Our model and code is available at github.com.

Visit

arxiv.org

Tasks

computer visionimage-text retrieval

Tags

Machine Learning

Similaires

Improving NER Tagging Performance in Low-Resource Languages via Multilingual LearningOptimal Transport Distillation for Robust Multilingual CLIP in Low-Resource Image RetrievalImproving Multilingual Semantic Textual Similarity with Shared Sentence Encoder for Low-resource LanguagesSimCSE Performance in Low-Resource Languages vs. Multilingual InfoNCE MethodsCLIP-Enhanced Annotation Projection for Cross-Lingual NER in Low-Resource LanguagesComparative Performance of Multilingual versus English Intermediate-Task Training on XTREME-R for Low-Resource Languages

Improving NER Tagging Performance in Low-Resource Languages via Multilingual Learning

Existing supervised solutions for Named Entity Recognition (NER) typically rely on a large annotated

Optimal Transport Distillation for Robust Multilingual CLIP in Low-Resource Image Retrieval

Benefiting from transformer-based pre-trained language models, neural ranking models have made signi

Improving Multilingual Semantic Textual Similarity with Shared Sentence Encoder for Low-resource Languages

Measuring the semantic similarity between two sentences (or Semantic Textual Similarity - STS) is fu

SimCSE Performance in Low-Resource Languages vs. Multilingual InfoNCE Methods

This report synthesises findings from 14 peer-reviewed papers addressing the following research ques

CLIP-Enhanced Annotation Projection for Cross-Lingual NER in Low-Resource Languages

Cross-lingual Named Entity Recognition (NER) leverages knowledge transfer between languages to ident

Comparative Performance of Multilingual versus English Intermediate-Task Training on XTREME-R for Low-Resource Languages

Intermediate-task training---fine-tuning a pretrained model on an intermediate task before fine-tuni