Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Homophobia and transphobia span identification in low-resource languages

Domaine:

natural language processing

Type de record:

dataset
Créateur:
KumKayPriBui
Éditeur:
UniUni
Éditeur:
Elsevier
Hôte:avatar
Online platforms have become prevalent because they promote free speech and group discussions. However, they also serve as platforms for hate speech, which can negatively impact the psychological well-being of vulnerable people. This is especially true for members of the LGBTQ+ community, who are often the targets of homophobia and transphobia in online environments. Our study makes three main contributions: (1) we developed a new dataset with span-level annotations for homophobia and transphobia in Tamil, English, and Marathi; (2) we employed advanced language models using BERT-based architectures, Conditional Random Field (CRF), and Bidirectional Long Short-Term Memory (BiLSTM) layers to enhance span-level detection of harmful content; and (3) we conducted benchmarking to evaluate the effectiveness of monolingual and multilingual models in detecting subtle forms of hate speech. The annotated dataset, which is collected from real-world social media (YouTube) content, provides diverse language contexts and enhances the representation of low-resource languages. The span-based detection approach enables models to detect subtle linguistic nuances, leading to more precise content moderation that accounts for cultural differences. The experimental results show that our models achieve effective span detection, which provides valuable information for creating inclusive moderation tools. Our research leads to the development of AI systems, and we aim to reduce the burden on moderators and improve the quality of online experiences for LGBTQ+ vulnerable.

Visit

doi.orgresearchrepository.universityofgalway.ie

Tasks

hate speech detectiontext classification

Tags

LGBTQ+ hate speechSpan-based classificationLow-resource languageBERT architectureSequence labeling

Licenses

CC BYCreative Commons Attribution 4.0 Internationalhttps://creativecommons.org/licenses/by/4.0/legalcode

Similaires

GlotLID: Language Identification for Low-Resource LanguagesLoanword Identification in Low-Resource Languages with Minimal Supervisionbitsa_nlp@LT-EDI-ACL2022: Leveraging Pretrained Language Models for Detecting Homophobia and Transphobia in Social Media CommentsSoumyaTeotia/Cross-Domain-Multi-Intent-Identification-in-Low-Resource-Languagesdev52003/Biased-News-Identification-For-Low-Resource-Languages-MalayalamAI and Low-Resource Languages

GlotLID: Language Identification for Low-Resource Languages

International audience Several recent papers have published good solutions for langua

Loanword Identification in Low-Resource Languages with Minimal Supervision

Bilingual resources play a very important role in many natural language processing tasks, especially

bitsa_nlp@LT-EDI-ACL2022: Leveraging Pretrained Language Models for Detecting Homophobia and Transphobia in Social Media Comments

Online social networks are ubiquitous and user-friendly. Nevertheless, it is vital to detect and mod

SoumyaTeotia/Cross-Domain-Multi-Intent-Identification-in-Low-Resource-Languages

Addressing the challenges of cross-domain and multi-intent classification in low-resource languages,

dev52003/Biased-News-Identification-For-Low-Resource-Languages-Malayalam

# 📰 Biased News Identification in Malayalam Media This repository contains the research work and to

AI and Low-Resource Languages

Artificial intelligence (AI) is rapidly transforming global communication, learning, and access to s