Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

MURAL: Multimodal, Multitask Representations Across Languages

Domain:

natural language processing

Record type:

model
Creator:
AssBalCheGuo
Publisher:
Und
Host:avatar
Both image-caption pairs and translation pairs provide the means to learn deep representations of and connections between languages. We use both types of pairs in MURAL (MUltimodal, MUltitask Representations Across Languages), a dual encoder that solves two tasks: 1) image-text matching and 2) translation pair matching. By incorporating billions of translation pairs, MURAL extends ALIGN (Jia et al.)--a state-of-the-art dual encoder learned from 1.8 billion noisy image-text pairs. When using the same encoders, MURAL's performance matches or exceeds ALIGN's cross-modal retrieval performance on well-resourced languages across several datasets. More importantly, it considerably improves performance on under-resourced languages, showing that text-text learning can overcome a paucity of image-caption examples for these languages. On the Wikipedia Image-Text dataset, for example, MURAL-base improves zero-shot mean recall by 8.1\% on average for eight under-resourced languages and by 6.8\% on average when fine-tuning. We additionally show that MURAL's text representations cluster not only with respect to genealogical connections but also based on areal linguistics, such as the Balkan Sprachbund.

Visit

doi.orgunderline.io

Tasks

image-text retrievalcomputer vision

Tags

Natural Language ProcessingQuestion-Answering SystemsSpeech ProcessingRobot VisionRobotics

Similar

Multimodal Multitask Representation Learning for Pathology Biobank Metadata PredictionThat Sounds Familiar: an Analysis of Phonetic Representations Transfer Across LanguagesAnalyzing LLMs' Knowledge Boundary Cognition Across Languages Through the Lens of Internal RepresentationsA Xhosa MuralMulti-Head Attention with Diversity for Learning Grounded Multilingual Multimodal RepresentationsLoraxBench: A Multitask, Multilingual Benchmark Suite for 20 Indonesian Languages

Multimodal Multitask Representation Learning for Pathology Biobank Metadata Prediction

Metadata are general characteristics of the data in a well-curated and condensed format, and have be

That Sounds Familiar: an Analysis of Phonetic Representations Transfer Across Languages

Only a handful of the world's languages are abundant with the resources that enable practical applic

Analyzing LLMs' Knowledge Boundary Cognition Across Languages Through the Lens of Internal Representations

While understanding the knowledge boundaries of LLMs is crucial to prevent hallucination, research o

A Xhosa Mural

Multi-Head Attention with Diversity for Learning Grounded Multilingual Multimodal Representations

With the aim of promoting and understanding the multilingual version of image search, we leverage vi

LoraxBench: A Multitask, Multilingual Benchmark Suite for 20 Indonesian Languages

As one of the world's most populous countries, with 700 languages spoken, Indonesia is behind in ter