Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Factors Influence Cross-Prompt Scoring of Arabic Essays

Domain:

natural language processing

Record type:

paper
Creator:
SulGow
Publisher:
Zenodo
Host:avatar
— Cross-prompt automated essay scoring (AES), which involves training on specific prompts and testing on unseen cases, represents a practical application scenario; however, this area remains underexplored for Arabic. Progress in AES is typically characterized by a succession of neural architectures that exhibit escalating levels of complexity. However, such progress is seldom evaluated against more basic choices such as tokenization, model-selection criteria, and text segmentation. We present the first controlled factorial study of cross-prompt Arabic AES on the LAILA dataset, crossing three architectures of increasing complexity (flat, hierarchical, multi-trait), three tokenizer–model configurations, and three random seeds, and benchmarking against a strong prompt-agnostic feature baseline. Our central finding is that tokenizer configuration drives performance far more than architecture: it outperforms each architectural step by approximately four times, covering a range of 0.16 to 0.19 in the overall Quadratic Weighted Kappa (QWK), the agreement metric used throughout. By contrast, the move from a flat to a hierarchical encoder yields only a small and seed-fragile gain (+0.045), multi-trait attention is mildly harmful (−0.02), and flat AraBERT matches the feature baseline (0.621 QWK) — which itself outperforms the hierarchical and multi-trait models in four of six configurations. We additionally provide evidence on how a seemingly straightforward "more-complex-is-better" result (M2 = 0.637 vs. M1 = 0.556 in our own early runs) is an artifact that dissolves once tokenization, selection criterion, and segmentation are each controlled. We release a corrected evaluation protocol — mean-per-prompt checkpoint selection and a sentence-based segmenter suited to Arabic — and argue that such controls are prerequisites for credible architecture claims in low-resource cross-prompt AES.

Visit

doi.org

Tasks

text classification

Tags

Automated Essay Scoring (AES), Arabic NLP, Cross-Prompt Essay Scoring, Tokenization Effects, AraBERT, Hierarchical Models, Multi-Trait Learning, Domain Generalization, Controlled Factorial Study, Evaluation Protocol

Licenses

Creative Commons Attribution 4.0 Internationalhttps://creativecommons.org/licenses/by/4.0/legalcode

Similar

Cross-Lingual Retrieval Augmented Prompt for Low-Resource LanguagesOn the Analysis of Cross-Lingual Prompt Tuning for Decoder-based Multilingual Modelchimaobim1/African-Chain-of-Thought-Prompt: African Context Prompt Template v1.0.0Two Waves of Berber Influence on Moroccan ArabicThe Arabic Influence on Northern BerberDesigning Annotation Guidelines for Trait-Based Arabic Automated Essay Scoring: A Systematic Methodology

Cross-Lingual Retrieval Augmented Prompt for Low-Resource Languages

Multilingual Pretrained Language Models (MPLMs) have shown their strong multilinguality in recent em

On the Analysis of Cross-Lingual Prompt Tuning for Decoder-based Multilingual Model

An exciting advancement in the field of multilingual models is the emergence of autoregressive model

chimaobim1/African-Chain-of-Thought-Prompt: African Context Prompt Template v1.0.0

First official release of the African Context Chain-of-Thought Prompt Engineering Template. Include

Two Waves of Berber Influence on Moroccan Arabic

The Arabic Influence on Northern Berber

Designing Annotation Guidelines for Trait-Based Arabic Automated Essay Scoring: A Systematic Methodology

Automated Essay Scoring (AES) fundamentally depends on high-quality annotated data, yet systematic a