Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

The Arabic AI Fingerprint: Stylometric Analysis and Detection of Large Language Models Text

Domain:

natural language processing

Record type:

papermodel
Creator:
Al-Ahm
Host:avatar
Large Language Models (LLMs) have achieved unprecedented capabilities in generating human-like text, posing subtle yet significant challenges for information integrity across critical domains, including education, social media, and academia, enabling sophisticated misinformation campaigns, compromising healthcare guidance, and facilitating targeted propaganda. This challenge becomes severe, particularly in under-explored and low-resource languages like Arabic. This paper presents a comprehensive investigation of Arabic machine-generated text, examining multiple generation strategies (generation from the title only, content-aware generation, and text refinement) across diverse model architectures (ALLaM, Jais, Llama, and GPT-4) in academic, and social media domains. Our stylometric analysis reveals distinctive linguistic patterns differentiating human-written from machine-generated Arabic text across these varied contexts. Despite their human-like qualities, we demonstrate that LLMs produce detectable signatures in their Arabic outputs, with domain-specific characteristics that vary significantly between different contexts. Based on these insights, we developed BERT-based detection models that achieved exceptional performance in formal contexts (up to 99.9\% F1-score) with strong precision across model architectures. Our cross-domain analysis confirms generalization challenges previously reported in the literature. To the best of our knowledge, this work represents the most comprehensive investigation of Arabic machine-generated text to date, uniquely combining multiple prompt generation methods, diverse model architectures, and in-depth stylometric analysis across varied textual domains, establishing a foundation for developing robust, linguistically-informed detection systems essential for preserving information integrity in Arabic-language contexts.

Visit

arxiv.org

Tasks

text classification

Tags

Computation and LanguageArtificial Intelligence

Similar

Arabic Large Language Models for Medical Text GenerationLarge Language Models for Arabic Sentiment Analysis and Dialect Detection: A Systematic ReviewAdversarial Evaluation of Large Language Models for Building Robust Offensive Language Detection in Moroccan ArabicCross-dialectal Arabic translation: comparative analysis on large language modelsA Survey of Large Language Models for Arabic Language and its DialectsEvaluation of Arabic Large Language Models on Moroccan Dialect

Arabic Large Language Models for Medical Text Generation

Efficient hospital management systems (HMS) are critical worldwide to address challenges such as ove

Large Language Models for Arabic Sentiment Analysis and Dialect Detection: A Systematic Review

## Overview and Motivation This research project is a systematic review that consolidates and criti

Adversarial Evaluation of Large Language Models for Building Robust Offensive Language Detection in Moroccan Arabic

Offensive language detection is crucial for ensuring safe and inclusive digital environments. Identi

Cross-dialectal Arabic translation: comparative analysis on large language models

Introduction Exploring Arabic dialects in Natural Language Processing (NLP) is essential to underst

A Survey of Large Language Models for Arabic Language and its Dialects

This survey offers a comprehensive overview of Large Language Models (LLMs) designed for Arabic lang

Evaluation of Arabic Large Language Models on Moroccan Dialect

Large Language Models (LLMs) have shown outstanding performance in many Natural Language Processing