Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Unsupervised Cross-Lingual Speech Emotion Recognition Using Pseudo Multilabel

Domaine:

natural language processing

Type de record:

papersoftware
Créateur:
Li,YanWan
Hôte:avatar
Speech Emotion Recognition (SER) in a single language has achieved remarkable results through deep learning approaches in the last decade. However, cross-lingual SER remains a challenge in real-world applications due to a great difference between the source and target domain distributions. To address this issue, we propose an unsupervised cross-lingual Neural Network with Pseudo Multilabel (NNPM) that is trained to learn the emotion similarities between source domain features inside an external memory adjusted to identify emotion in cross-lingual databases. NNPM introduces a novel approach that leverages external memory to store source domain features and generates pseudo multilabel for each target domain data by computing the similarities between the external memory and the target domain features. We evaluate our approach on multiple different languages of speech emotion databases. Experimental results show our proposed approach significantly improves the weighted accuracy (WA) across multiple low-resource languages on Urdu, Skropus, ShEMO, and EMO-DB corpus. To facilitate further research, code is available at github.com

Visit

arxiv.org

Tasks

emotion identificationspeech processing

Tags

Audio and Speech Processing

Similaires

Unsupervised ASR via Cross-Lingual Pseudo-LabelingEffectiveness of text to speech pseudo labels for forced alignment and cross lingual pretrained models for low resource speech recognitionUnsupervised Speech RecognitionCross-Corpus Multilingual Speech Emotion Recognition: Amharic vs. Other LanguagesUnsupervised Cross-Lingual Part-of-Speech Tagging with Monolingual Corpora OnlyYeshimebetBayu/Amharic-MultiLabel-Emotion-Dataset

Unsupervised ASR via Cross-Lingual Pseudo-Labeling

Recent work has shown that it is possible to train an $\textit{unsupervised}$ automatic speech recog

Effectiveness of text to speech pseudo labels for forced alignment and cross lingual pretrained models for low resource speech recognition

In the recent years end to end (E2E) automatic speech recognition (ASR) systems have achieved promis

Unsupervised Speech Recognition

Despite rapid progress in the recent past, current speech recognition systems still require labeled

Cross-Corpus Multilingual Speech Emotion Recognition: Amharic vs. Other Languages

In a conventional Speech emotion recognition (SER) task, a classifier for a given language is traine

Unsupervised Cross-Lingual Part-of-Speech Tagging with Monolingual Corpora Only

Due to the scarcity of part-of-speech annotated data, existing studies on low-resource languages typ

YeshimebetBayu/Amharic-MultiLabel-Emotion-Dataset

A multi-label emotion classification dataset for Amharic language, designed for research and develop