Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

MELLA: Bridging Linguistic Capability and Cultural Groundedness for Low-Resource Language MLLMs

Domaine:

natural language processing

Type de record:

paperdatasetmodel
Créateur:
GaoFeiCheChe
Hôte:avatar
Multimodal Large Language Models (MLLMs) have shown remarkable performance in high-resource languages. However, their effectiveness diminishes significantly in the contexts of low-resource languages. Current multilingual enhancement methods are often limited to text modality or rely solely on machine translation. While such approaches help models acquire basic linguistic capabilities and produce "thin descriptions", they neglect the importance of multimodal informativeness and cultural groundedness, both of which are crucial for serving low-resource language users effectively. To bridge this gap, in this study, we identify two significant objectives for a truly effective MLLM in low-resource language settings, namely 1) linguistic capability and 2) cultural groundedness, placing special emphasis on cultural awareness. To achieve these dual objectives, we propose a dual-source strategy that guides the collection of data tailored to each goal, sourcing native web alt-text for culture and MLLM-generated captions for linguistics. As a concrete implementation, we introduce MELLA, a multimodal, multilingual dataset. Experiment results show that after fine-tuning on MELLA, there is a general performance improvement for the eight languages on various MLLM backbones, with models producing "thick descriptions". We verify that the performance gains are from both cultural knowledge enhancement and linguistic capability enhancement. Our dataset can be found at opendatalab.com.

Visit

arxiv.org

Tasks

computer visionimage-text retrieval

Tags

Computer Vision and Pattern RecognitionArtificial IntelligenceComputation and Language

Similaires

Harnessing Linguistic Dissimilarity for Language Generalization on Unseen Low-Resource VarietiesBridging the Gap: Enhancing LLM Performance for Low-Resource African Languages with New Benchmarks, Fine-Tuning, and Cultural AdjustmentsA Forensic Linguistic Dataset for Offensive Content Detection in Low-Resource Language: SetswanaBridging context gaps in low-resource language chatbots through multilevel attention and hybrid embedding approachesLinguistic Diversity Metrics in Intermediate-Task Selection for Low-Resource Language Zero-Shot AccuracyBridging the Language Gap in Text-to-SQL: Adapting LLMs for Chichewa in a Low-Resource Setting

Harnessing Linguistic Dissimilarity for Language Generalization on Unseen Low-Resource Varieties

Low-resource language varieties used by specific groups remain neglected in the development of Multi

Bridging the Gap: Enhancing LLM Performance for Low-Resource African Languages with New Benchmarks, Fine-Tuning, and Cultural Adjustments

Large Language Models (LLMs) have shown remarkable performance across various tasks, yet significant

A Forensic Linguistic Dataset for Offensive Content Detection in Low-Resource Language: Setswana

Developing Monolingual Setswana Datasets for Offensive Content Detection Reproducibility Package, Me

Bridging context gaps in low-resource language chatbots through multilevel attention and hybrid embedding approaches

Conversational agents for low-resource languages (LRLs), such as Igbo, face major challenges, includ

Linguistic Diversity Metrics in Intermediate-Task Selection for Low-Resource Language Zero-Shot Accuracy

Intermediate-task training---fine-tuning a pretrained model on an intermediate task before fine-tuni

Bridging the Language Gap in Text-to-SQL: Adapting LLMs for Chichewa in a Low-Resource Setting