Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

SLAyiNG: A Diverse and Community-validated Dataset of Queer Slang

Domaine:

natural language processing

Type de record:

dataset
Créateur:
VelHirZidWic
Éditeur:
arXiv
Hôte:avatar
Queer vernacular is rarely studied in NLP, despite advancements in resources and evaluation for other sociolects and informal language. Because of this, NLP systems often process queer language incorrectly, e.g., they misclassify it as hate speech or generate negative responses. To address this problem, we propose Slaying, the first real-world dataset of English queer slang. Slaying is community-validated, and includes over 500 queer slang terms that pertain to more than 20 queer subcommunities. We argue that queer language data resources have great potential in NLP -- e.g., as components of large pretraining corpora and as the basis for benchmarks -- and can improve queer users' experience of NLP systems. We leverage Slaying for two novel findings in support of this argument: (i) For a number of language models, we show that they are unbiased towards the queer community, but at the same time unable to process its language, i.e., absence of representation bias does not entail the absence of linguistic bias. (ii) Model performance on queer slang varies across queer subcommunities; it is generally worse for slang pertaining to African-American and Latine communities. These findings are relevant for both the queer NLP and the broader ML communities. Slaying is available to the public, and open to future revisions and extensions. Warning: This paper contains profane and potentially offensive language. Preprint

Visit

doi.org

Tasks

text classification

Tags

Computation and Language (cs.CL)FOS: Computer and information sciences

Licenses

Creative Commons Attribution 4.0 Internationalhttps://creativecommons.org/licenses/by/4.0/legalcode

Similaires

jumadesigns/Naija-Street-Slang-Dataset-NSSD-ProjectGeographically Diverse DatasetAdvancing Conversational AI with Shona Slang: A Dataset and Hybrid Model for Digital InclusionFusha–Darija Evaluation Dataset: Human-Validated Phase-2 ExpansionPalm: A Culturally Inclusive and Linguistically Diverse Dataset for Arabic LLMsQueer studies and religion in Southern Africa: The production of queer Christian subjects

jumadesigns/Naija-Street-Slang-Dataset-NSSD-Project

Welcome to the Naija Street Slang Dataset (NSSD)! An open-source initiative for collecting and curat

Geographically Diverse Dataset

Geographically diverse dataset introduced in the paper Dataset Diversity: Measuring and Mitigating G

Advancing Conversational AI with Shona Slang: A Dataset and Hybrid Model for Digital Inclusion

African languages remain underrepresented in natural language processing (NLP), with most corpora li

Fusha–Darija Evaluation Dataset: Human-Validated Phase-2 Expansion

Human-validated Phase-2 expansion of the Fusha–Darija evaluation project. The release contains publi

Palm: A Culturally Inclusive and Linguistically Diverse Dataset for Arabic LLMs

As large language models (LLMs) become increasingly integrated into daily life, ensuring their cultu

Queer studies and religion in Southern Africa: The production of queer Christian subjects

Abstract The question of how to write about queer Africa has been a significant