Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Detecting Machine-Generated Arabic Text Using AraBERT and LSTM: Toward Trustworthy NLP in Low-Resource Languages

Domain:

natural language processing

Record type:

paper
Creator:
BarIbrAl
Publisher:
Zenodo
Host:avatar
Deepfake text generation has emerged as a serious challenge in the age of advanced language models, particularly in low-resource languages like Arabic. This study presents a deep learning-based approach to detect synthetic Arabic text generated by AI systems. We propose a binary classification framework combining AraBERT embeddings with a Long Short-Term Memory (LSTM) network. A balanced dataset of 87,452 samples was constructed using real Arabic text and synthetic text generated via AraGPT2. Our best-performing model achieved a test accuracy of 99.5%, demonstrating strong generalization and detection capability. This work contributes to enhancing Arabic NLP security and offers a foundation for future multilingual deepfake detection systems.

Visit

doi.orgzenodo.org

Tasks

text classification

Licenses

Creative Commons Attribution 4.0 Internationalhttps://creativecommons.org/licenses/by/4.0/legalcodeCopyright© [2025] Tarek Barhoum et al. This is an open access preprint distributed under the Creative Commons Attribution 4.0 International License (CC BY 4.0), which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited.http://rightsstatements.org/vocab/InC/1.0/