Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

mSTEB: Massively Multilingual Evaluation of LLMs on Speech and Text Tasks

Domaine:

natural language processing

Type de record:

paperdataset
Créateur:
BeyVerMa,Ala
Hôte:avatar
Large Language models (LLMs) have demonstrated impressive performance on a wide range of tasks, including in multimodal settings such as speech. However, their evaluation is often limited to English and a few high-resource languages. For low-resource languages, there is no standardized evaluation benchmark. In this paper, we address this gap by introducing mSTEB, a new benchmark to evaluate the performance of LLMs on a wide range of tasks covering language identification, text classification, question answering, and translation tasks on both speech and text modalities. We evaluated the performance of leading LLMs such as Gemini 2.0 Flash and GPT-4o (Audio) and state-of-the-art open models such as Qwen 2 Audio and Gemma 3 27B. Our evaluation shows a wide gap in performance between high-resource and low-resource languages, especially for languages spoken in Africa and Americas/Oceania. Our findings show that more investment is needed to address their under-representation in LLMs coverage. Accepted to ASRU 2025

Visit

arxiv.org

Tasks

language identification

Tags

Computation and LanguageMachine LearningSoundAudio and Speech Processing

Similaires

Omnilingual SONAR: Cross-Lingual and Cross-Modal Sentence Embeddings Bridging Massively Multilingual Text and SpeechCommon Voice: A Massively-Multilingual Speech CorpusMassively Multilingual Text Translation For Low-Resource LanguagesBabel-ImageNet: Massively Multilingual Evaluation of Vision-and-Language RepresentationsCS-FLEURS: A Massively Multilingual and Code-Switched Speech DatasetAfriVox: Probing Multilingual and Accent Robustness of Speech LLMs

Omnilingual SONAR: Cross-Lingual and Cross-Modal Sentence Embeddings Bridging Massively Multilingual Text and Speech

Cross-lingual sentence encoders typically cover only a few hundred languages and often trade downstr

Common Voice: A Massively-Multilingual Speech Corpus

The Common Voice corpus is a massively-multilingual collection of transcribed speech intended for speech technology research and development. Common Voice is designed for Automatic Speech Recognition purposes but can be useful in other domains (e.g. language identi

Massively Multilingual Text Translation For Low-Resource Languages

Translation into severely low-resource languages has both the cultural goal of saving and reviving t

Babel-ImageNet: Massively Multilingual Evaluation of Vision-and-Language Representations

Vision-and-language (VL) models with separate encoders for each modality (e.g., CLIP) have become the go-to models for zero-shot image classification and image-text retrieval. The bulk of the evaluation of these models is, however, performed with English text only:

CS-FLEURS: A Massively Multilingual and Code-Switched Speech Dataset

We present CS-FLEURS, a new dataset for developing and evaluating code-switched speech recognition a

AfriVox: Probing Multilingual and Accent Robustness of Speech LLMs

Recent advances in multimodal and speech-native large language models (LLMs) have delivered impressi