Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

mBLIP: Efficient Bootstrapping of Multilingual Vision-LLMs

Domain:

natural language processing

Record type:

papermodel
Creator:
GeiJaiTimGla
Host:avatar
Modular vision-language models (Vision-LLMs) align pretrained image encoders with (frozen) large language models (LLMs) and post-hoc condition LLMs to `understand' the image input. With the abundance of readily available high-quality English image-text data as well as strong monolingual English LLMs, the research focus has been on English-only Vision-LLMs. Multilingual vision-language models are still predominantly obtained via expensive end-to-end pretraining, resulting in comparatively smaller models, trained on limited multilingual image data supplemented with text-only multilingual corpora. We present mBLIP, the first Vision-LLM leveraging multilingual LLMs, which we obtain in a computationally efficient manner on consumer-level hardware. To this end, we \textit{re-align} an image encoder previously tuned to an English LLM to a new, multilingual LLM using only a few million multilingual training examples derived from a mix of vision-and-language tasks, which we obtain by machine-translating high-quality English data to 95 languages. On the IGLUE benchmark and XM3600, mBLIP yields results competitive with state-of-the-art models and it greatly outperforms strong English-only Vision-LLMs like Llava 1.5. We release our model, code, and train data at \url{github.com. ALVR Workshop 2024

Visit

arxiv.org

Tags

Computer Vision and Pattern RecognitionComputation and Language

Similar

LLMs Beyond English: Scaling the Multilingual Capability of LLMs with Cross-Lingual FeedbackControlling Language Confusion in Multilingual LLMsMultilingual jailbreaking of LLMs using low-resource languagesSpokenNativQA: Multilingual Everyday Spoken Queries for LLMsAfriVox: Probing Multilingual and Accent Robustness of Speech LLMsMultimodal Classification System for Hausa Using LLMs and Vision Transformers

LLMs Beyond English: Scaling the Multilingual Capability of LLMs with Cross-Lingual Feedback

To democratize large language models (LLMs) to most natural languages, it is imperative to make thes

Controlling Language Confusion in Multilingual LLMs

Large language models often suffer from language confusion, a phenomenon in which responses are part

Multilingual jailbreaking of LLMs using low-resource languages

Large Language Models (LLMs) remain vulnerable to jailbreak attempts that circumvent safety guardrai

SpokenNativQA: Multilingual Everyday Spoken Queries for LLMs

Large Language Models (LLMs) have demonstrated remarkable performance across various disciplines and

AfriVox: Probing Multilingual and Accent Robustness of Speech LLMs

Recent advances in multimodal and speech-native large language models (LLMs) have delivered impressi

Multimodal Classification System for Hausa Using LLMs and Vision Transformers

This paper presents a classification-based Vi-sual Question Answering (VQA) system for theHausa lang