Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Multimodal Depression Detection from Speech and Text Using a Fusion Neural Network

Domain:

healthcare

Record type:

paper
Creator:
KinYebKamCal
Publisher:
RSI
Host:
Depression is among the leading causes of disability worldwide, yet its detection continues to rely heavily on subjective clinical interviews and self-report instruments that are difficult to scale, particularly in low-resource regions. This paper presents the design, implementation, and empirical validation of a multimodal fusion neural network for automatic depression detection from speech and text, developed with specific attention to closing the near-total absence of African research contributions in this rapidly growing area of computing. The proposed system extracts spectral-prosodic acoustic features from voice recordings using a convolutional encoder, and contextual semantic features from transcribed text using a transformer-based language encoder, before an attention gate learns to weight the relative contribution of each modality per instance ahead of a fully connected classifier. The architecture was implemented and validated end-to-end as a feasibility study on two accessible, weak-label proxy corpora: the RAVDESS acted-emotion speech dataset, with sad and calm recordings relabeled as a depression-like class, and the dair-ai/emotion text corpus, with sadness and fear posts relabeled likewise. On held-out test data, the fusion model reached 90.28% accuracy and 0.9609 AUC-ROC, and, most notably, recovered the positive-class recall that the audio-only model lost almost entirely (21.05% versus 84.21%), demonstrating that the attention-gated fusion mechanism functions as designed. This paper reports the background, problem definition, related work, the complete mathematical formulation of the implemented pipeline, the empirical results of this proxy validation, and a discussion of what they do and do not establish, closing with concrete recommendations for advancing the work toward a clinically meaningful, Africa-relevant screening tool.

Visit

doi.org

Tasks

speech processing

Similar

Enhancing Biometric Security Using Artificial Neural Network-Based Multimodal Fusion of Facial Recognition and Fingerprint IdentificationA part of speech tagger for Yoruba language text using deep neural networkAutomated Amharic Hate Speech Posts and Comments Detection Model Using Recurrent Neural NetworkMultimodal Amharic Hate Speech Detection Using Deep LearningImage–Text Multimodal Sentiment Analysis Framework of Assamese News Articles Using Late FusionMalaria detection using Deep Convolution Neural Network

Enhancing Biometric Security Using Artificial Neural Network-Based Multimodal Fusion of Facial Recognition and Fingerprint Identification

International audience Biometric authentication systems based on a single modality re

A part of speech tagger for Yoruba language text using deep neural network

Automated Amharic Hate Speech Posts and Comments Detection Model Using Recurrent Neural Network

Abstract During the last few years, social activities over the internet especially on soci

Multimodal Amharic Hate Speech Detection Using Deep Learning

Image–Text Multimodal Sentiment Analysis Framework of Assamese News Articles Using Late Fusion

Before the arrival of the web as a corpus, people detected positive and negative news based on the u

Malaria detection using Deep Convolution Neural Network

The latest WHO report showed that the number of malaria cases climbed to 219 million last year, two