Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Fleurs-SLU: A Massively Multilingual Benchmark for Spoken Language Understanding

Domain:

natural language processing

Record type:

paperdataset
Creator:
SchVulGlaAde
Host:avatar
Spoken language understanding (SLU) is indispensable for half of all living languages that lack a formal writing system. Unlike for high-resource languages, for these languages, we cannot offload semantic understanding of speech to the cascade of automatic speech recognition (ASR) and text-based large language models (LLMs). Even if low-resource languages possess a writing system, ASR for these languages remains unreliable due to limited bimodal speech and text training data. Nonetheless, the evaluation of multilingual SLU is limited to shallow tasks such as intent classification or language identification. This is why we present Fleurs-SLU, a multilingual SLU benchmark that encompasses (i) 692 hours of speech for topical utterance classification in 102 languages and (ii) multiple-choice question answering via listening comprehension spanning 944 hours of speech across 92 languages. We extensively evaluate end-to-end speech classification models, cascaded systems that combine speech-to-text transcription with subsequent LLM-based classification, and multimodal speech-LLMs on Fleurs-SLU. Our results show that cascaded systems are more robust in multilingual SLU, though well-pretrained speech encoders can perform competitively in topical speech classification. Closed-source speech-LLMs match or surpass the performance of cascaded systems. We observe a strong correlation between robust multilingual ASR, effective speech-to-text translation, and strong multilingual SLU, indicating mutual benefits between acoustic and semantic speech representations.

Visit

arxiv.org

Tasks

speech processing

Tags

Computation and LanguageArtificial Intelligence

Similar

TARIC-SLU: A Tunisian Benchmark Dataset for Spoken Language UnderstandingMVL-SIB: A Massively Multilingual Vision-Language Benchmark for Cross-Modal Topical MatchingCS-FLEURS: A Massively Multilingual and Code-Switched Speech DatasetXTREME: A Massively Multilingual Multi-task Benchmark for Evaluating Cross-lingual GeneralizationMultiBLiMP 1.0: A Massively Multilingual Benchmark of Linguistic Minimal PairsSERENGETI: Massively Multilingual Language Models for Africa

TARIC-SLU: A Tunisian Benchmark Dataset for Spoken Language Understanding

MVL-SIB: A Massively Multilingual Vision-Language Benchmark for Cross-Modal Topical Matching

Existing multilingual vision-language (VL) benchmarks often only cover a handful of languages. Conse

CS-FLEURS: A Massively Multilingual and Code-Switched Speech Dataset

We present CS-FLEURS, a new dataset for developing and evaluating code-switched speech recognition a

XTREME: A Massively Multilingual Multi-task Benchmark for Evaluating Cross-lingual Generalization

The Cross-lingual Natural Language Inference (XNLI) corpus is a crowd-sourced collection of 5,000 test and 2,500 dev pairs for the MultiNLI corpus. The pairs are annotated with textual entailment and translated into 14 languages: French, Spanish, German, Greek, Bu

MultiBLiMP 1.0: A Massively Multilingual Benchmark of Linguistic Minimal Pairs

We introduce MultiBLiMP 1.0, a massively multilingual benchmark of linguistic minimal pairs, coverin

SERENGETI: Massively Multilingual Language Models for Africa

Multilingual language models (MLMs) acquire valuable, generalizable linguistic information during pretraining and have advanced the state of the art on task-specific finetuning. So far, only ~ 28 out of ~2,000 African languages are covered in existing language mode