Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Fair-OBNC: Correcting Label Noise for Fairer Datasets

Domain:

socioeconomic

Record type:

paper
Creator:
OliJesFerSal
Host:avatar

Data used by automated decision-making systems, such as Machine Learning models, often reflects discriminatory behavior that occurred in the past. These biases in the training data are sometimes related to label noise, such as in COMPAS, where more African-American offenders are wrongly labeled as having a higher risk of recidivism when compared to their White counterparts. Models trained on such biased data may perpetuate or even aggravate the biases with respect to sensitive information, such as gender, race, or age. However, while multiple label noise correction approaches are available in the literature, these focus on model performance exclusively. In this work, we propose Fair-OBNC, a label noise correction method with fairness considerations, to produce training datasets with measurable demographic parity. The presented method adapts Ordering-Based Noise Correction, with an adjusted criterion of ordering, based both on the margin of error of an ensemble, and the potential increase in the observed demographic parity of the dataset. We evaluate Fair-OBNC against other different pre-processing techniques, under different scenarios of controlled label noise. Our results show that the proposed method is the overall better alternative within the pool of label correction methods, being capable of attaining better reconstructions of the original labels. Models trained in the corrected data have an increase, on average, of 150% in demographic parity, when compared to models trained in data with noisy labels, across the considered levels of label noise.

Visit

doi.org

Licenses

info:eu-repo/semantics/openAccessCreative Commons Attribution 4.0 Internationalhttps://creativecommons.org/licenses/by/4.0/legalcode

Similar

Correcting the Tamazight Portions of FLORES+ and OLDI Seed DatasetsCross-lingual NER robustness to label noise under unlabeled data volume variationMulti-source Teacher-Student Learning for Cross-lingual NER in Low-Resource Languages with Label NoiseSegmenting Subtitles for Correcting ASR Segmentation ErrorsDo Fairer Elections Increase the Responsiveness of Politicians?Correcting FLORES Evaluation Dataset for Four African Languages

Correcting the Tamazight Portions of FLORES+ and OLDI Seed Datasets

Cross-lingual NER robustness to label noise under unlabeled data volume variation

Cross-lingual Named Entity Recognition (NER) leverages knowledge transfer between languages to ident

Multi-source Teacher-Student Learning for Cross-lingual NER in Low-Resource Languages with Label Noise

Cross-lingual transfer learning enables NLP for low-resource languages by leveraging labeled data fr

Segmenting Subtitles for Correcting ASR Segmentation Errors

Typical ASR systems segment the input audio into utterances using purely acoustic information, which

Do Fairer Elections Increase the Responsiveness of Politicians?

Leveraging novel experimental designs and 2,160 months of Constituency Development Fund (CDF) spendi

Correcting FLORES Evaluation Dataset for Four African Languages

This paper describes the corrections made to the FLORES evaluation (dev and devtest) dataset for fou