Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

MASSIVE: A 1M-Example Multilingual Natural Language Understanding Dataset with 51 Typologically-Diverse Languages

Domaine:

natural language processing

Type de record:

paper
We present the MASSIVE dataset--Multilingual Amazon Slu resource package (SLURP) for Slot-filling, Intent classification, and Virtual assistant Evaluation. MASSIVE contains 1M realistic, parallel, labeled virtual assistant utterances spanning 51 languages, 18 domains, 60 intents, and 55 slots. MASSIVE was created by tasking professional translators to localize the English-only SLURP dataset into 50 typologically diverse languages from 29 genera. We also present modeling results on XLM-R and mT5, including exact match accuracy, intent classification accuracy, and slot-filling F1 score. We have released our dataset, modeling code, and models publicly.

Visit

arxiv.orgwww.amazon.science

Connected records

dataset

Tasks

machine translation

Languages

AfrikaansAmharicSwahili

Licenses

https://github.com/alexa/massive/blob/main/LICENSE.txt

Similaires

MULTI3NLU++: A Multilingual, Multi-Intent, Multi-Domain Dataset for Natural Language Understanding in Task-Oriented DialogueNatural Language Understanding Datasets for African LanguagesMasakhaPOS: Part-of-Speech Tagging for Typologically Diverse African LanguagesAimeM250/Natural-Language-Understanding-System-for-African-LanguagesTyDi QA: A Benchmark for Information-Seeking Question Answering in Typologically Diverse LanguagesThe Fragility of Chain-of-Thought Monitoring Across Typologically Diverse Languages

MULTI3NLU++: A Multilingual, Multi-Intent, Multi-Domain Dataset for Natural Language Understanding in Task-Oriented Dialogue

Task-oriented dialogue (TOD) systems have been widely deployed in many industries as they deliver mo

Natural Language Understanding Datasets for African Languages

Natural Language Understanding (NLU)

Natural Language Understanding (NLU) is a fundamental building block of goal-oriented dialogue systems like Alexa, Siri, Google Assistant, and Cortana (Fig 1). One of the major challenges of NLU is predicting the use

MasakhaPOS: Part-of-Speech Tagging for Typologically Diverse African Languages

In this paper, we present MasakhaPOS, the largest part-of-speech (POS) dataset for 20 typologically

AimeM250/Natural-Language-Understanding-System-for-African-Languages

The following repository contains a Natural Language Understanding for African languages (NLU) proje

TyDi QA: A Benchmark for Information-Seeking Question Answering in Typologically Diverse Languages

Confidently making progress on multilingual modeling requires challenging, trustworthy evaluations. We present TyDi QA---a question answering dataset covering 11 typologically diverse languages with 204K question-answer pairs. The languages of TyDi QA are diverse w

The Fragility of Chain-of-Thought Monitoring Across Typologically Diverse Languages

Chain-of-thought (CoT) monitoring has been proposed as a promising safety mechanism for detecting mi