Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

MENASpeechBank: A Reference Voice Bank with Persona-Conditioned Multi-Turn Conversations for AudioLLMs

Domaine:

natural language processing

Type de record:

paperdataset
Créateur:
AliBhaNanCho
Hôte:avatar
Audio large language models (AudioLLMs) enable instruction-following over speech and general audio, but progress is increasingly limited by the lack of diverse, conversational, instruction-aligned speech-text data. This bottleneck is especially acute for persona-grounded interactions and dialectal coverage, where collecting and releasing real multi-speaker recordings is costly and slow. We introduce MENASpeechBank, a reference speech bank comprising about 18K high-quality utterances from 124 speakers spanning multiple MENA countries, covering English, Modern Standard Arabic (MSA), and regional Arabic varieties. Building on this resource, we develop a controllable synthetic data pipeline that: (i) constructs persona profiles enriched with World Values Survey-inspired attributes, (ii) defines a taxonomy of about 5K conversational scenarios, (iii) matches personas to scenarios via semantic similarity, (iv) generates about 417K role-play conversations with an LLM where the user speaks as the persona and the assistant behaves as a helpful agent, and (v) synthesizes the user turns by conditioning on reference speaker audio to preserve speaker identity and diversity. We evaluate both synthetic and human-recorded conversations and provide detailed analysis. We will release MENASpeechBank and the generated conversations publicly for the community. Foundation Models, Large Language Models, Native, Speech Models, Arabic, AI-persona, Persona-conditioned-conversations

Visit

arxiv.org

Tasks

speech processing

Tags

SoundArtificial IntelligenceComputation and LanguageAudio and Speech Processing68T50F.2.2; I.2.7

Similaires

FineTome Single Turn Conversations - AmharicMulti-Task Instruction Tuning via Data Scheduling for Low-Resource Arabic AudioLLMsSLANG-GraphRAG: Multi-Layered Retrieval with Domain-Specific Knowledge for Low Resource Social Media Conversationsjabelopitso/voice-bank-africa-lingoConversations with a Xhosa Child: Dialogue in Early ChildhoodAn Inter-lingual Reference Approach For Multi-Lingual Ontology Matching

FineTome Single Turn Conversations - Amharic

This dataset contains 83,290 conversational examples translated from English to Amharic, providing h

Multi-Task Instruction Tuning via Data Scheduling for Low-Resource Arabic AudioLLMs

Audio large language models (LLMs) enable unified speech understanding and generation, but adapting

SLANG-GraphRAG: Multi-Layered Retrieval with Domain-Specific Knowledge for Low Resource Social Media Conversations

Emotion classification on social media is especially difficult when texts include informal, cultural

jabelopitso/voice-bank-africa-lingo

VoiceBank Africa Project Overview VoiceBank Africa is a revolutionary voice-activated banking web a

Conversations with a Xhosa Child: Dialogue in Early Childhood

Pre-dialogue, proto-dialogue and dialogue are described and examples of these in the language develo

An Inter-lingual Reference Approach For Multi-Lingual Ontology Matching

Ontologies are considered as the backbone of the Semantic Web. With the rising success of the Semant