Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

ALHD: A Large-Scale and Multigenre Benchmark Dataset for Arabic LLM-Generated Text Detection

Domain:

natural language processing

Record type:

paperdataset
Creator:
KhaZub
Host:avatar
We introduce ALHD, the first large-scale comprehensive Arabic dataset explicitly designed to distinguish between human- and LLM-generated texts. ALHD spans three genres (news, social media, reviews), covering both MSA and dialectal Arabic, and contains over 400K balanced samples generated by three leading LLMs and originated from multiple human sources, which enables studying generalizability in Arabic LLM-genearted text detection. We provide rigorous preprocessing, rich annotations, and standardized balanced splits to support reproducibility. In addition, we present, analyze and discuss benchmark experiments using our new dataset, in turn identifying gaps and proposing future research directions. Benchmarking across traditional classifiers, BERT-based models, and LLMs (zero-shot and few-shot) demonstrates that fine-tuned BERT models achieve competitive performance, outperforming LLM-based models. Results are however not always consistent, as we observe challenges when generalizing across genres; indeed, models struggle to generalize when they need to deal with unseen patterns in cross-genre settings, and these challenges are particularly prominent when dealing with news articles, where LLM-generated texts resemble human texts in style, which opens up avenues for future research. ALHD establishes a foundation for research related to Arabic LLM-detection and mitigating risks of misinformation, academic dishonesty, and cyber threats. 47 pages, 15 figures. Dataset available at Zenodo: doi.org Codebase available at GitHub: github.com

Visit

arxiv.org

Tasks

text classification

Tags

Computation and LanguageArtificial IntelligenceMachine Learning

Similar

SARD: A Large-Scale Synthetic Arabic OCR Dataset for Book-Style Text RecognitionNigerian Academic Writing Corpus: Pre-AI Benchmark for AI-Generated Text DetectionDesign and Evaluation of a Parallel Classifier for Large-Scale Arabic TextFraud Detection Using Large-scale Imbalance DatasetBenchmark Dataset for DiaLex, A Benchmark for Evaluating Multidialectal Arabic Word EmbeddingsBenchmark Dataset for Amharic Text Summarization

SARD: A Large-Scale Synthetic Arabic OCR Dataset for Book-Style Text Recognition

Arabic Optical Character Recognition (OCR) is essential for converting vast amounts of Arabic print

Nigerian Academic Writing Corpus: Pre-AI Benchmark for AI-Generated Text Detection

A curated corpus of pre-AI era (2005–2022) Nigerian academic writing paired with AI-generated equiva

Design and Evaluation of a Parallel Classifier for Large-Scale Arabic Text

Fraud Detection Using Large-scale Imbalance Dataset

In the context of machine learning, an imbalanced classification problem states to a dataset in whic

Benchmark Dataset for DiaLex, A Benchmark for Evaluating Multidialectal Arabic Word Embeddings

Word embeddings are a core component of modern natural language processing systems, making the ability to thoroughly evaluate them a vital task. We describe DiaLex, a benchmark for intrinsic evaluation of dialectal Arabic word embeddings. DiaLex covers five importa

Benchmark Dataset for Amharic Text Summarization

Standardizing Amharic NLP: A Benchmark Dataset for Amharic Text Summarization