Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

SpokenNativQA: Multilingual Everyday Spoken Queries for LLMs

Domaine:

natural language processing

Type de record:

paperdataset
Créateur:
AlaHasCho
Hôte:avatar
Large Language Models (LLMs) have demonstrated remarkable performance across various disciplines and tasks. However, benchmarking their capabilities with multilingual spoken queries remains largely unexplored. In this study, we introduce SpokenNativQA, the first multilingual and culturally aligned spoken question-answering (SQA) dataset designed to evaluate LLMs in real-world conversational settings. The dataset comprises approximately 33,000 naturally spoken questions and answers in multiple languages, including low-resource and dialect-rich languages, providing a robust benchmark for assessing LLM performance in speech-based interactions. SpokenNativQA addresses the limitations of text-based QA datasets by incorporating speech variability, accents, and linguistic diversity. We benchmark different ASR systems and LLMs for SQA and present our findings. We released the data at (huggingface.co) and the experimental scripts at (llmebench.qcri.org) for the research community. Spoken Question Answering, Multilingual LLMs, Speech-based Evaluation, Dialectal Speech, Low-resource Languages, Multimodal Benchmarking, Conversational AI, Speech-to-Text QA, Real-world Interaction, Natural Language Understanding

Visit

arxiv.org

Tasks

automatic speech recognitionquestion answeringspeech processing

Tags

Computation and LanguageArtificial Intelligence68T50I.2.7

Similaires

BerkGuny/Sequential-fine-tuning-for-improving-everyday-cultural-knowledge-in-multilingual-LLMsBLEnD: A Benchmark for LLMs on Everyday Knowledge in Diverse Cultures and LanguagesCross-Lingual Auto Evaluation for Assessing Multilingual LLMsMultilingual Spoken Words CorpusControlling Language Confusion in Multilingual LLMsWASIL: In-the-Wild Arabic Spoken Interactions with LLMs

BerkGuny/Sequential-fine-tuning-for-improving-everyday-cultural-knowledge-in-multilingual-LLMs

Developed a two-stage multilingual LLM fine-tuning pipeline that improves culturally grounded questi

BLEnD: A Benchmark for LLMs on Everyday Knowledge in Diverse Cultures and Languages

Large language models (LLMs) often lack culture-specific knowledge of daily life, especially across diverse regions and non-English languages. Existing benchmarks for evaluating LLMs' cultural sensitivities are limited to a single language or collected from onli

Cross-Lingual Auto Evaluation for Assessing Multilingual LLMs

Evaluating machine-generated text remains a significant challenge in NLP, especially for non-English

Multilingual Spoken Words Corpus

Multilingual Spoken Words Corpus is a large and growing audio dataset of spoken words in 50 languages collectively spoken by over 5 billion people, for academic research and commercial applications in keyword spotting and spoken term search, licensed under CC-BY 4.

Controlling Language Confusion in Multilingual LLMs

Large language models often suffer from language confusion, a phenomenon in which responses are part

WASIL: In-the-Wild Arabic Spoken Interactions with LLMs

Large Language Models (LLMs) voice assistants are commonly built as cascaded Automatic Speech recogn