Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Mohithavelagapudi/Audio-Driven-Hate-Speech-Detection-in-Telugu

Domaine:

natural language processing

Type de record:

datasetmodel
Créateur:
Moh
Hôte:
Low-resource multimodal hate speech detection leveraging acoustic and textual representations for robust moderation in Telugu. # 🎧 Audio-Driven Hate Speech Detection in Telugu **Low-resource multimodal hate speech detection leveraging acoustic and textual representations for robust moderation in Telugu.** ---- ### 🚀 Overview While hate speech detection has progressed rapidly for English, Telugu — with over 83 million speakers — still lacks annotated resources. This project introduces the first multimodal Telugu hate speech dataset and a suite of audio-, text-, and fusion-based models for comprehensive detection. ### 🧠 Core Highlights - 🗂️ First Telugu hate-speech dataset (2 hours of annotated audio–text pairs). - 🔊 Multimodal pipeline integrating acoustic and textual cues. - ⚙️ Evaluated OpenSMILE, Wav2Vec2, LaBSE, and XLM-R baselines. - 🎯 Achieved 91 % accuracy (audio) and 89 % (text); fusion improved robustness. ---- ## 🧩 Abstract This study fills a critical resource gap in Telugu hate-speech detection. A manually annotated 2-hour multimodal dataset was curated from YouTube. Acoustic (OpenSMILE + SVM) and textual (LaBSE) models achieved 91 % and 89 % accuracy, respectively. Fusion approaches highlight the complementary role of vocal prosody and linguistic cues. ---- ## 🎯 Problem Statement | Challenge | Description | | ------------------------ | ----------------------------------------------------------------------------- | | 🗣️ **Low-Resource Gap** | Telugu lacks labeled corpora and pretrained models for hate-speech detection. | | 🔊 **Modality Gap** | Text-only systems ignore vocal signals (tone, sarcasm, aggression). | 💡 Goal: Develop a multimodal framework combining speech and text for richer, context-aware classification. ---- ### 📊 Dataset: DravLangGuard | Attribute | Description | | ----------------------------- | -------------------------------------- | | **Source** | YouTube (≥ 50 K subscribers) …

Visit

github.com

Tasks

hate speech detectionspeech processingtext classification

Tags

binaryclassificationearly-fusionlabselate-fusionlibrosam-bertmlpmulticlass-classificationmultimodal-learningopensmile+5