Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Separating the Wheat from the Chaff with BREAD: An open-source benchmark and metrics to detect redundancy in text

Domaine:

natural language processing

Type de record:

paperdataset
Créateur:
CasWanPap
Hôte:avatar
Data quality is a problem that perpetually resurfaces throughout the field of NLP, regardless of task, domain, or architecture, and remains especially severe for lower-resource languages. A typical and insidious issue, affecting both training data and model output, is data that is repetitive and dominated by linguistically uninteresting boilerplate, such as price catalogs or computer-generated log files. Though this problem permeates many web-scraped corpora, there has yet to be a benchmark to test against, or a systematic study to find simple metrics that generalize across languages and agree with human judgements of data quality. In the present work, we create and release BREAD, a human-labeled benchmark on repetitive boilerplate vs. plausible linguistic content, spanning 360 languages. We release several baseline CRED (Character REDundancy) scores along with it, and evaluate their effectiveness on BREAD. We hope that the community will use this resource to develop better filtering methods, and that our reference implementations of CRED scores can become standard corpus evaluation tools, driving the development of cleaner language modeling corpora, especially in low-resource languages. Accepted to GEM workshop 2023; 6 pages

Visit

arxiv.org

Tasks

text classification

Tags

Computation and LanguageMachine Learning

Similaires

Separating Grains from the Chaff: Using Data Filtering to Improve Multilingual Translation for Low-Resourced African LanguagesAmgyna/Africa-text-Open-Source-Dataset-Response of bread wheat to sulfur and phosphorus fertilizers in the north central EthiopiaHow to Evaluate Speech Translation with Source-Aware Neural MT MetricsSeedling and Adult Plant Resistance in the Ethiopian Bread Wheat Landraces to Stripe Rust DiseaseImprovement in yield of bread wheat cultivars released in Ethiopia from 1949 to 1987

Separating Grains from the Chaff: Using Data Filtering to Improve Multilingual Translation for Low-Resourced African Languages

We participated in the WMT 2022 Large-Scale Machine Translation Evaluation for the African Languages

Amgyna/Africa-text-Open-Source-Dataset-

# Africa-text-Open-Source-Dataset- ## 🇹🇿Maelezo ya Kiswahili Karibu kwenye **African Text Datasets*

Response of bread wheat to sulfur and phosphorus fertilizers in the north central Ethiopia

Abstract Background Emerging research evidences since few years

How to Evaluate Speech Translation with Source-Aware Neural MT Metrics

Automatic evaluation of ST systems is typically performed by comparing translation hypotheses with o

Seedling and Adult Plant Resistance in the Ethiopian Bread Wheat Landraces to Stripe Rust Disease

High yielding farmers’ bread wheat cultivars are threatened by emerging race(s) of stripe

Improvement in yield of bread wheat cultivars released in Ethiopia from 1949 to 1987