Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

BLUFF: Benchmarking the Detection of False and Synthetic Content across 58 Low-Resource Languages

Domain:

natural language processing

Record type:

paperdataset
Creator:
LucMurUchAl-
Host:avatar
Multilingual falsehoods threaten information integrity worldwide, yet detection benchmarks remain confined to English or a few high-resource languages, leaving low-resource linguistic communities without robust defense tools. We introduce BLUFF, a comprehensive benchmark for detecting false and synthetic content, spanning 79 languages with over 202K samples, combining human-written fact-checked content (122K+ samples across 57 languages) and LLM-generated content (79K+ samples across 71 languages). BLUFF uniquely covers both high-resource "big-head" (20) and low-resource "long-tail" (59) languages, addressing critical gaps in multilingual research on detecting false and synthetic content. Our dataset features four content types (human-written, LLM-generated, LLM-translated, and hybrid human-LLM text), bidirectional translation (English$\leftrightarrow$X), 39 textual modification techniques (36 manipulation tactics for fake news, 3 AI-editing strategies for real news), and varying edit intensities generated using 19 diverse LLMs. We present AXL-CoI (Adversarial Cross-Lingual Agentic Chainof-Interactions), a novel multi-agentic framework for controlled fake/real news generation, paired with mPURIFY, a quality filtering pipeline ensuring dataset integrity. Experiments reveal state-of-theart detectors suffer up to 25.3% F1 degradation on low-resource versus high-resource languages. BLUFF provides the research community with a multilingual benchmark, extensive linguistic-oriented benchmark evaluation, comprehensive documentation, and opensource tools to advance equitable falsehood detection. Dataset and code are available at: jsl5710.github.io

Visit

arxiv.org

Tags

Computation and Language

Similar

Robustness of Synthetic vs. Human-Annotated Grammatical Error Detection Models in Low-Resource LanguagesGrammatical Error Detection Performance in Low-Resource Languages: Zero-Shot Synthetic vs. Human-Annotated BaselinesVLURes: Benchmarking VLM Visual and Linguistic Understanding in Low-Resource LanguagesBenchmarking Neural and Statistical Machine Translation on Low-Resource African LanguagesToxicity Red-Teaming: Benchmarking LLM Safety in Singapore's Low-Resource Languagessrshettyy/anote-low-resource-benchmarking

Robustness of Synthetic vs. Human-Annotated Grammatical Error Detection Models in Low-Resource Languages

Grammatical Error Detection (GED) methods rely heavily on human annotated error corpora. However, th

Grammatical Error Detection Performance in Low-Resource Languages: Zero-Shot Synthetic vs. Human-Annotated Baselines

Grammatical Error Detection (GED) methods rely heavily on human annotated error corpora. However, th

VLURes: Benchmarking VLM Visual and Linguistic Understanding in Low-Resource Languages

Vision Language Models (VLMs) are pivotal for advancing perception in intelligent agents. Yet, evalu

Benchmarking Neural and Statistical Machine Translation on Low-Resource African Languages

Research in machine translation (MT) is developing at a rapid pace. However, most work in the community has focused on languages where large amounts of digital resources are available. In this study, we benchmark state of the art statistical and neural machine tran

Toxicity Red-Teaming: Benchmarking LLM Safety in Singapore's Low-Resource Languages

The advancement of Large Language Models (LLMs) has transformed natural language processing; however

srshettyy/anote-low-resource-benchmarking

A cross-lingual active learning pipeline evaluating few-shot prompt optimizations on low-resource la