Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

PolySpeech-100: A Large-Scale Benchmark for Speech Understanding Across 100+ Languages and Dialects

Domain:

natural language processing

Record type:

paperdataset
Creator:
YanRuaWu,Liu, Yu
Host:avatar
While End-to-End (E2E) Speech-Large Language Models (Speech-LLMs) are rapidly evolving, their evaluation methodologies remain limited to the era of simple transcription. Existing benchmarks suffer from three critical limitations: a pronounced bias towards high-resource languages, a focus on low-level recognition (ASR) rather than semantic reasoning, and a neglect of regional dialects. To bridge this gap, we introduce PolySpeech-100, a massive-scale benchmark designed to assess `native-level' speech comprehension across 110 linguistic variants. We employ a novel hybrid construction pipeline that augments gold-standard human recordings with instruction-driven synthetic speech, allowing us to cover 19 distinct Chinese dialects and over 80 low-resource languages. Extensive evaluation of 22 state-of-the-art models (including Gemini-3, GPT-Audio, and Qwen2.5-Omni) yields pivotal insights. First, we demonstrate that open-source E2E models outperform Cascade (ASR+LLM) systems on heavy dialects, proving that direct audio processing preserves critical paralinguistic cues and prosodic features (e.g., intonation, stress) that are often lost in standard transcription. Second, we reveal a significant performance gap: while commercial models maintain robustness, open-source models suffer catastrophic degradation on low-resource languages. Finally, counter-intuitively, we observe that under standard zero-shot settings, Chain-of-Thought prompting frequently degrades speech understanding performance for most evaluated models, revealing a potential modality alignment gap in current architectures. PolySpeech-100 establishes a rigorous standard for the next generation of inclusive, omni-capable Speech-LLMs. The data, demo, and code are publicly available at github.com. 19 pages, 13 figures, KDD 2026

Visit

arxiv.org

Tasks

automatic speech recognitionspeech processing

Tags

Computation and LanguageArtificial IntelligenceAudio and Speech Processing

Similar

Global PIQA: Evaluating Commonsense Reasoning Across 100+ Languages and CulturesGlobal PIQA: Evaluating Physical Commonsense Reasoning Across 100+ Languages and Cultures v0.1Script collapse in multilingual ASR: A reference-free metric and 100-pair benchmarkSaptak: A Large-scale Multi-Regional Benchmark Dataset for Poly-Dialectal Neural Machine Translation between Standard Bangla and Regional Dialects, and among Regional DialectsMangaUB: A Manga Understanding Benchmark for Large Multimodal ModelsLet's Play Across Cultures: A Large Multilingual, Multicultural Benchmark for Assessing Language Models' Understanding of Sports

Global PIQA: Evaluating Commonsense Reasoning Across 100+ Languages and Cultures

To date, there exist almost no culturally-specific evaluation benchmarks for large language models (

Global PIQA: Evaluating Physical Commonsense Reasoning Across 100+ Languages and Cultures v0.1

To date, there exist almost no culturally-specific evaluation benchmarks for large language models (LLMs) that cover a large number of languages and cultures. In this paper, we present Global PIQA, a participatory commonsense reasoning benchmark for over 100 lan

Script collapse in multilingual ASR: A reference-free metric and 100-pair benchmark

Word error rate (WER) is the dominant metric for automatic speech recognition, yet it cannot detect

Saptak: A Large-scale Multi-Regional Benchmark Dataset for Poly-Dialectal Neural Machine Translation between Standard Bangla and Regional Dialects, and among Regional Dialects

This dataset is a comprehensive parallel corpus developed for poly-dialectal neural machine translat

MangaUB: A Manga Understanding Benchmark for Large Multimodal Models

Manga is a popular medium that combines stylized drawings and text to convey stories. As manga panel

Let's Play Across Cultures: A Large Multilingual, Multicultural Benchmark for Assessing Language Models' Understanding of Sports

Language Models (LMs) are primarily evaluated on globally popular sports, often overlooking regional