Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

MELLA: Bridging Linguistic Capability and Cultural Groundedness for Low-Resource Language MLLMs

Domain:

natural language processing

Record type:

paperdatasetmodel
Creator:
GaoFeiCheChe
Host:avatar
Multimodal Large Language Models (MLLMs) have shown remarkable performance in high-resource languages. However, their effectiveness diminishes significantly in the contexts of low-resource languages. Current multilingual enhancement methods are often limited to text modality or rely solely on machine translation. While such approaches help models acquire basic linguistic capabilities and produce "thin descriptions", they neglect the importance of multimodal informativeness and cultural groundedness, both of which are crucial for serving low-resource language users effectively. To bridge this gap, in this study, we identify two significant objectives for a truly effective MLLM in low-resource language settings, namely 1) linguistic capability and 2) cultural groundedness, placing special emphasis on cultural awareness. To achieve these dual objectives, we propose a dual-source strategy that guides the collection of data tailored to each goal, sourcing native web alt-text for culture and MLLM-generated captions for linguistics. As a concrete implementation, we introduce MELLA, a multimodal, multilingual dataset. Experiment results show that after fine-tuning on MELLA, there is a general performance improvement for the eight languages on various MLLM backbones, with models producing "thick descriptions". We verify that the performance gains are from both cultural knowledge enhancement and linguistic capability enhancement. Our dataset can be found at opendatalab.com.

Visit

arxiv.org

Tasks

computer visionimage-text retrieval

Tags

Computer Vision and Pattern RecognitionArtificial IntelligenceComputation and Language

Similar

Harnessing Linguistic Dissimilarity for Language Generalization on Unseen Low-Resource VarietiesBridging the Gap: Enhancing LLM Performance for Low-Resource African Languages with New Benchmarks, Fine-Tuning, and Cultural AdjustmentsA Forensic Linguistic Dataset for Offensive Content Detection in Low-Resource Language: SetswanaBridging context gaps in low-resource language chatbots through multilevel attention and hybrid embedding approachesLinguistic Diversity Metrics in Intermediate-Task Selection for Low-Resource Language Zero-Shot AccuracyBridging the Language Gap in Text-to-SQL: Adapting LLMs for Chichewa in a Low-Resource Setting

Harnessing Linguistic Dissimilarity for Language Generalization on Unseen Low-Resource Varieties

Low-resource language varieties used by specific groups remain neglected in the development of Multi

Bridging the Gap: Enhancing LLM Performance for Low-Resource African Languages with New Benchmarks, Fine-Tuning, and Cultural Adjustments

Large Language Models (LLMs) have shown remarkable performance across various tasks, yet significant

A Forensic Linguistic Dataset for Offensive Content Detection in Low-Resource Language: Setswana

Developing Monolingual Setswana Datasets for Offensive Content Detection Reproducibility Package, Me

Bridging context gaps in low-resource language chatbots through multilevel attention and hybrid embedding approaches

Conversational agents for low-resource languages (LRLs), such as Igbo, face major challenges, includ

Linguistic Diversity Metrics in Intermediate-Task Selection for Low-Resource Language Zero-Shot Accuracy

Intermediate-task training---fine-tuning a pretrained model on an intermediate task before fine-tuni

Bridging the Language Gap in Text-to-SQL: Adapting LLMs for Chichewa in a Low-Resource Setting