Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Characterizing Types of Convolution in Deep Convolutional Recurrent Neural Networks for Robust Speech Emotion Recognition

Domaine:

natural language processing

Type de record:

paper
Créateur:
HuaNar
Hôte:avatar
Deep convolutional neural networks are being actively investigated in a wide range of speech and audio processing applications including speech recognition, audio event detection and computational paralinguistics, owing to their ability to reduce factors of variations, for learning from speech. However, studies have suggested to favor a certain type of convolutional operations when building a deep convolutional neural network for speech applications although there has been promising results using different types of convolutional operations. In this work, we study four types of convolutional operations on different input features for speech emotion recognition under noisy and clean conditions in order to derive a comprehensive understanding. Since affective behavioral information has been shown to reflect temporally varying of mental state and convolutional operation are applied locally in time, all deep neural networks share a deep recurrent sub-network architecture for further temporal modeling. We present detailed quantitative module-wise performance analysis to gain insights into information flows within the proposed architectures. In particular, we demonstrate the interplay of affective information and the other irrelevant information during the progression from one module to another. Finally we show that all of our deep neural networks provide state-of-the-art performance on the eNTERFACE'05 corpus. Revised Submission to IEEE Transactions

Visit

arxiv.org

Tasks

emotion identificationspeech processing

Tags

Machine LearningComputation and LanguageMultimediaSound

Similaires

Language Identification Using Deep Convolutional Recurrent Neural NetworksExplainable deep convolutional neural networks for insect pest recognitionThe convolutional neural networks for Amazigh speech recognition systemTowards Robust Arabic Speech Emotion Recognition with Deep LearningConvolutional neural network for speech emotion recognition in the Moroccan Arabic dialect languageExploring data augmentation for Amazigh speech recognition with convolutional neural networks

Language Identification Using Deep Convolutional Recurrent Neural Networks

Language Identification (LID) systems are used to classify the spoken language from a given audio sa

Explainable deep convolutional neural networks for insect pest recognition

International audience Fungal infestation of crops is critical to food security as it

The convolutional neural networks for Amazigh speech recognition system

Towards Robust Arabic Speech Emotion Recognition with Deep Learning

Speech Emotion Recognition (SER) aims to identify a speaker's emotional state from audio signals. Wh

Convolutional neural network for speech emotion recognition in the Moroccan Arabic dialect language

Extracting the speaker's emotional state has become an active research topic lately due to the deman

Exploring data augmentation for Amazigh speech recognition with convolutional neural networks