A comprehensive framework for detecting hate speech in the low-resource Telugu language using multimodal (audio and text) analysis
# 🎧 Audio-Driven Hate Speech Detection in Telugu
**Low-resource multimodal hate speech detection leveraging acoustic and textual representations for robust moderation in Telugu.**
----
### 🚀 Overview
While hate speech detection has progressed rapidly for English, Telugu — with over 83 million speakers — still lacks annotated resources.
This project introduces the first multimodal Telugu hate speech dataset and a suite of audio-, text-, and fusion-based models for comprehensive detection.
### 🧠 Core Highlights
- 🗂️ First Telugu hate-speech dataset (2 hours of annotated audio–text pairs).
- 🔊 Multimodal pipeline integrating acoustic and textual cues.
- ⚙️ Evaluated OpenSMILE, Wav2Vec2, LaBSE, and XLM-R baselines.
- 🎯 Achieved 91 % accuracy (audio) and 89 % (text); fusion improved robustness.
----
## 🧩 Abstract
This study fills a critical resource gap in Telugu hate-speech detection.
A manually annotated 2-hour multimodal dataset was curated from YouTube.
Acoustic (OpenSMILE + SVM) and textual (LaBSE) models achieved 91 % and 89 % accuracy, respectively.
Fusion approaches highlight the complementary role of vocal prosody and linguistic cues.
----
## 🎯 Problem Statement
| Challenge | Description |
| ------------------------ | ----------------------------------------------------------------------------- |
| 🗣️ **Low-Resource Gap** | Telugu lacks labeled corpora and pretrained models for hate-speech detection. |
| 🔊 **Modality Gap** | Text-only systems ignore vocal signals (tone, sarcasm, aggression). |
💡 Goal: Develop a multimodal framework combining speech and text for richer, context-aware classification.
----
### 📊 Dataset: DravLangGuard
| Attribute | Description |
| ----------------------------- | -------------------------------------- |
| **Source** | YouTube (≥ 50 K subscribers) …