
Abstract
Background: Large Language Models (LLMs), particularly ChatGPT-4o, have shown promise as automated translation quality assessment tools, yet their reliability against established human annotation frameworks remains empirically underexplored especially for Arabic, a morphologically complex language underrepresented in LLM evaluation research. Methods: Drawing on an existing corpus of 270 student translation errors from a Saudi university CAT course (25 texts; Arabic–English, bi-directional), this study positions ChatGPT-4o as a prospective automated quality annotator and subjects it to systematic empirical validation against expert human judgement within the Multidimensional Quality Metrics (MQM) framework — an approach that, to the author’s knowledge, has not previously been applied to Arabic–English student translation data. A purpose-built, standardised prompt protocol generated structured JSON error outputs from ChatGPT for each source–translation pair. Agreement between ChatGPT and the human gold standard was then quantified through Cohen’s Kappa, computed separately for error category and severity classification. Three widely-used automated evaluation metrics — BLEU, COMET, and BERTScore — were correlated with human-annotated error counts to examine their predictive validity in this educational context. Results: ChatGPT achieved substantial agreement with human raters on error category classification (κ = 0.648, 73.0% agreement) and moderate agreement on severity assessment (κ = 0.478, 66.7%). Agreement was highest for Locale Convention (κ = 0.724) and lowest for Accuracy errors (κ = 0.312). Translation direction (AR→EN vs EN→AR) did not significantly moderate agreement rates (χ² = 0.132, p = .716). All automated metrics showed strong negative correlations with human-annotated error counts (r = −.962 to −.982). Conclusions: ChatGPT-4o offers reliable first-pass diagnostic classification of Arabic–English student translation errors at the category level. However, its severity judgement remains in the moderate range, indicating that human expert assessment is not yet replaceable for fine-grained quality decisions. The study provides a replicable MQM prompt protocol with direct applications in translation pedagogy and automated feedback systems.
Keywords: machine translation evaluation; MQM; ChatGPT; Arabic–English translation; inter-rater reliability; natural language processing; translation pedagogy