Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

UniVerse: Benchmarking and Enhancing LALMs on Culturally Inclusive Low-Resource Music Understanding

Domaine:

natural language processing

Type de record:

paper
Créateur:
ZhoWu,Xu,Zhe
Éditeur:
arXiv
Hôte:avatar
Recent advances in large audio-language models (LALMs) have significantly improved performance in tasks such as music captioning, genre classification, and sound event detection. However, limited attention has been paid to improving their adaptability across diverse musical traditions, particularly folk music rooted in distinct cultural contexts. Folk-music traditions are typically resource-scarce, unevenly represented across regions, and poorly documented. Even when such samples appear in large-scale pre-training, LALMs often fail to capture their structural and stylistic characteristics, partly due to the absence of dedicated evaluation protocols and training solutions. To address these limitations, we introduce UniVerse, a reproducible solution for low-resource music understanding. Specifically, we propose UniVerseBench, a benchmark of 5,042 Q&A pairs across more than 38 cultural and linguistic entities, constructed via an expert-guided yet highly automated pipeline. In parallel, we construct a fully automated, model-generated multi-turn dialogue training dataset UniVerseSet. By training LALMs on UniVerseSet, we systematically adapt and investigate representative multimodal imbalance learning strategies across both dense and Mixture-of-Experts (MoE) architectures. Experimental results indicate that fully automated data curation combined with imbalance-aware training yields non-trivial improvements, but models still struggle to capture fine-grained acoustic features, indicating a gap between surface-level alignment and deep musical comprehension. 21 pages, 7 figures, 8 tables

Visit

doi.org

Tags

Sound (cs.SD)Multimedia (cs.MM)FOS: Computer and information sciences

Licenses

Creative Commons Attribution 4.0 Internationalhttps://creativecommons.org/licenses/by/4.0/legalcode

Similaires

VLURes: Benchmarking VLM Visual and Linguistic Understanding in Low-Resource LanguagesCulturally-Grounded Chain-of-Thought (CG-CoT):Enhancing LLM Performance on Culturally-Specific Tasks in Low-Resource LanguagesFrom Understanding to Appreciating Music Cross- Culturallysrshettyy/anote-low-resource-benchmarkingBenchmarking Neural and Statistical Machine Translation on Low-Resource African LanguagesBenchmarking of Low-Resource Machine Translation Systems

VLURes: Benchmarking VLM Visual and Linguistic Understanding in Low-Resource Languages

Vision Language Models (VLMs) are pivotal for advancing perception in intelligent agents. Yet, evalu

Culturally-Grounded Chain-of-Thought (CG-CoT):Enhancing LLM Performance on Culturally-Specific Tasks in Low-Resource Languages

Large Language Models (LLMs) struggle with culturally-specific reasoning tasks, particularly in low-

From Understanding to Appreciating Music Cross- Culturally

It has long been debated which aspects of music perception are universal and which are developed onl

srshettyy/anote-low-resource-benchmarking

A cross-lingual active learning pipeline evaluating few-shot prompt optimizations on low-resource la

Benchmarking Neural and Statistical Machine Translation on Low-Resource African Languages

Research in machine translation (MT) is developing at a rapid pace. However, most work in the community has focused on languages where large amounts of digital resources are available. In this study, we benchmark state of the art statistical and neural machine tran

Benchmarking of Low-Resource Machine Translation Systems

Assessing the performance of machine translation systems is of critical value, especially to languag