Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Bloom Library: Multimodal Datasets in 300+ Languages for a Variety of Downstream Tasks

Domain:

natural language processing

Record type:

paperdataset
Creator:
LeoNemManFil
Host:avatar
We present Bloom Library, a linguistically diverse set of multimodal and multilingual datasets for language modeling, image captioning, visual storytelling, and speech synthesis/recognition. These datasets represent either the most, or among the most, multilingual datasets for each of the included downstream tasks. In total, the initial release of the Bloom Library datasets covers 363 languages across 32 language families. We train downstream task models for various languages represented in the data, showing the viability of the data for future work in low-resource, multimodal NLP and establishing the first known baselines for these downstream tasks in certain languages (e.g., Bisu [bzi], with an estimated population of 700 users). Some of these first-of-their-kind baselines are comparable to state-of-the-art performance for higher-resourced languages. The Bloom Library datasets are released under Creative Commons licenses on the Hugging Face datasets hub to catalyze more linguistically diverse research in the included downstream tasks. 14 pages, 1 figure, 3 tables, accepted to and presented at EMNLP 2022

Visit

arxiv.org

Tags

Computation and LanguageArtificial Intelligence

Similar

Effects of Swahili Monolingual Tokenizer on Downstream TasksMultimodal Pretraining as Intermediate Tasks for Zero-Shot Cross-Lingual Transfer in Low-Resource Languages on XTREME-RAfri Code Datasets (A collection of datasets for code generation in African languages)Integration of TLI in Multilingual Pre-training for Robust Cross-lingual Transfer in African Multimodal TasksLarge Multimodal Models for Low-Resource Languages: A SurveyAfriInstruct: Instruction Tuning of African Languages for Diverse Tasks

Effects of Swahili Monolingual Tokenizer on Downstream Tasks

Multimodal Pretraining as Intermediate Tasks for Zero-Shot Cross-Lingual Transfer in Low-Resource Languages on XTREME-R

Intermediate-task training---fine-tuning a pretrained model on an intermediate task before fine-tuni

Afri Code Datasets (A collection of datasets for code generation in African languages)

Training and evaluating Large Language Models (LLMs) for code generation, building AI-powered coding

Integration of TLI in Multilingual Pre-training for Robust Cross-lingual Transfer in African Multimodal Tasks

This paper studies zero-shot cross-lingual transfer of vision-language models. Specifically, we focu

Large Multimodal Models for Low-Resource Languages: A Survey

In this survey, we systematically analyze techniques used to adapt large multimodal models (LMMs) fo

AfriInstruct: Instruction Tuning of African Languages for Diverse Tasks

Large language models (LLMs) for African languages perform worse compared to their performance in hi