Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

WoNBias: A Dataset for Classifying Bias & Prejudice Against Women in Bengali Text

Domain:

natural language processing

Record type:

dataset
Creator:
AssIslTaf
Publisher:
Und
Host:avatar
This paper presents WoNBias, a curated Bengali dataset to identify gender-based biases, stereotypes, and harmful language directed at women. It merges digital sources- social media, blogs, news- with offline tactics comprising surveys and focus groups, alongside some existing corpora to compile a total of 31,484 entries (10,656 negative; 10,170 positive; 10,658 neutral). WoNBias reflects the sociocultural subtleties of bias in both Bengali digital and offline conversations. By bridging online and offline biased contexts, the dataset supports content moderation, policy interventions, and equitable NLP research for Bengali, a low-resource language critically underserved by existing tools. WoNBias aims to combat systemic gender discrimination against women on digital platforms, empowering researchers and practitioners to combat harmful narratives in Bengali-speaking communities.

Visit

doi.orgunderline.io

Tasks

hate speech detectiontext classification

Tags

Computational LinguisticsNatural Language ProcessingArtificial Intelligence

Similar

BanglaFakeNews: A Curated Dataset for Bengali Fake News DetectionPatiGonit22K: A Comprehensive Dataset for Solving Complex Bengali MWPsPolitically Aggressive Bengali DatasetA Comparative Analysis of Policy Frameworks in Overcoming Institutional Bias Against Women-Led Sustainable EnterprisesSSC-BanglaTutor: A Curriculum-Aligned Bengali Dataset for Intelligent Tutoring SystemsBengali Idiom Detection: A BIO-Annotated Dataset for Figurative Language Processing

BanglaFakeNews: A Curated Dataset for Bengali Fake News Detection

BanglaFakeNews is a large-scale, curated dataset developed for fake news detection in the Bengali la

PatiGonit22K: A Comprehensive Dataset for Solving Complex Bengali MWPs

Mathematical Word Problems (MWPs) are an important benchmark for evaluating natural language underst

Politically Aggressive Bengali Dataset

The Politically Aggressive Bengali Dataset (PABD) is a manually annotated dataset developed to facil

A Comparative Analysis of Policy Frameworks in Overcoming Institutional Bias Against Women-Led Sustainable Enterprises

This study analytically compared the effectiveness of policy intervention frameworks in addressing i

SSC-BanglaTutor: A Curriculum-Aligned Bengali Dataset for Intelligent Tutoring Systems

This dataset comprises a Bengali-language educational corpus specifically curated to support the fin

Bengali Idiom Detection: A BIO-Annotated Dataset for Figurative Language Processing

This dataset includes a corpus of the Bengali language for idiom identification and sequence tagging