Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

"RoU-AudioTox: Roman Urdu Audio-Text Dataset for Offensive Speech Detection and Transcription"

Domaine:

natural language processing

Type de record:

dataset
Créateur:
MarAsmAlmHaf
Éditeur:
IEE
Hôte:avatar
"Since online communications are growing, those who speak Roman Urdu experience more problems with offensive statements and online bullying. English gets a lot of attention from researchers and is supported with many specialized datasets, but Roman Urdu does not, despite being used on social media for Urdu communication. Our solution to bridge this gap is a Roman Urdu offensive multimodal dataset that has been labeled as Abusive, Threatening, Anti-National or Neutral. Our dataset includes recorded audio from a variety of speakers, along with well-labeled transcripts. This database allows researchers to test models for turning offensive speech into text and for detecting offensive content, mainly for low-resource language Roman Urdu. As a result, many of us in the community benefit from being able to fight cyberbullying and communicate safely and easily online."

Visit

doi.orgieee-dataport.org

Tasks

automatic speech recognitionhate speech detectionspeech processingtext classification

Licenses

Creative Commons Attribution 4.0 Internationalhttps://creativecommons.org/licenses/by/4.0/legalcode

Similaires

Hate Speech Detection in Roman UrduFine-Tuning Large Language Models with QLoRA for Offensive Language Detection in Roman Urdu-English Code-Mixed TextRoman Urdu Phishing SMS Dataset for Smishing Detection ResearchCombining FastText and Glove Word Embedding for Offensive and Hate speech Text DetectionAn Attention Based Neural Network for Code Switching Detection: English & Roman UrduSequence to Sequence Networks for Roman-Urdu to Urdu Transliteration

Hate Speech Detection in Roman Urdu

Hate speech is a specific type of controversial content that is widely legislated as a crime that mu

Fine-Tuning Large Language Models with QLoRA for Offensive Language Detection in Roman Urdu-English Code-Mixed Text

The use of derogatory terms in languages that employ code mixing, such as Roman Urdu, presents chall

Roman Urdu Phishing SMS Dataset for Smishing Detection Research

This dataset contains labeled Roman Urdu/English SMS messages for phishing (smishing) detection rese

Combining FastText and Glove Word Embedding for Offensive and Hate speech Text Detection

Combining FastText and Glove Word Embedding for Offensive and Hate speech Text Detection

Poster presented at the Deep Learning Indaba 2022 by Nabil BADRI

An Attention Based Neural Network for Code Switching Detection: English & Roman Urdu

Code-switching is a common phenomenon among people with diverse lingual background and is widely use

Sequence to Sequence Networks for Roman-Urdu to Urdu Transliteration

Neural Machine Translation models have replaced the conventional phrase based statistical translatio