Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Computational Error Profiling in Student Arabic–English Translation: How Reliably Can ChatGPT Diagnose and Classify Errors Against Human MQM Assessment?

Domain:

natural language processing

Record type:

paper
Creator:
Naj
Publisher:
Zenodo
Host:avatar

Abstract

Background: Large Language Models (LLMs), particularly ChatGPT-4o, have shown promise as automated translation quality assessment tools, yet their reliability against established human annotation frameworks remains empirically underexplored especially for Arabic, a morphologically complex language underrepresented in LLM evaluation research. Methods: Drawing on an existing corpus of 270 student translation errors from a Saudi university CAT course (25 texts; Arabic–English, bi-directional), this study positions ChatGPT-4o as a prospective automated quality annotator and subjects it to systematic empirical validation against expert human judgement within the Multidimensional Quality Metrics (MQM) framework — an approach that, to the author’s knowledge, has not previously been applied to Arabic–English student translation data. A purpose-built, standardised prompt protocol generated structured JSON error outputs from ChatGPT for each source–translation pair. Agreement between ChatGPT and the human gold standard was then quantified through Cohen’s Kappa, computed separately for error category and severity classification. Three widely-used automated evaluation metrics — BLEU, COMET, and BERTScore — were correlated with human-annotated error counts to examine their predictive validity in this educational context. Results: ChatGPT achieved substantial agreement with human raters on error category classification (κ = 0.648, 73.0% agreement) and moderate agreement on severity assessment (κ = 0.478, 66.7%). Agreement was highest for Locale Convention (κ = 0.724) and lowest for Accuracy errors (κ = 0.312). Translation direction (AR→EN vs EN→AR) did not significantly moderate agreement rates (χ² = 0.132, p = .716). All automated metrics showed strong negative correlations with human-annotated error counts (r = −.962 to −.982). Conclusions: ChatGPT-4o offers reliable first-pass diagnostic classification of Arabic–English student translation errors at the category level. However, its severity judgement remains in the moderate range, indicating that human expert assessment is not yet replaceable for fine-grained quality decisions. The study provides a replicable MQM prompt protocol with direct applications in translation pedagogy and automated feedback systems.

 

Keywords: machine translation evaluation; MQM;  ChatGPT;  Arabic–English translation; inter-rater reliability; natural language processing; translation pedagogy

Visit

doi.org

Tasks

machine translation

Languages

Ndasa

Licenses

info:eu-repo/semantics/openAccessCreative Commons Attribution 4.0 Internationalhttps://creativecommons.org/licenses/by/4.0/legalcode

Similar

ChatGPT for Arabic Grammatical Error CorrectionInvestigating Calquing In Arabic How Translation Can Shape The Future Of ArabicHuman VS. AI Translation Accuracy: A Comparative Study of English-Arabic legal Contract TranslationComputational Approaches to Arabic-English Code-SwitchingERROR ANALYSIS OF TIGRINYA– ENGLISH MACHINE  TRANSLATION SYSTEMSChatGPT for Arabic-English Translation: Evaluating the Accuracy ChatGPT للترجمة من العربية إلى الإنجليزية: تقييم الدقة ChatGPT pour la traduction arabe-anglais : évaluation de l'exactitude ChatGPT para traducción árabe-inglés: evaluación de la precisión

ChatGPT for Arabic Grammatical Error Correction

Recently, large language models (LLMs) fine-tuned to follow human instruction have exhibited signifi

Investigating Calquing In Arabic How Translation Can Shape The Future Of Arabic

This paper investigates the influx of Anglicism and Frenchism into the Arabic lexicon. The focus wil

Human VS. AI Translation Accuracy: A Comparative Study of English-Arabic legal Contract Translation

As the title indicates, this study aims to conduct a comparative analysis to evaluate the a

Computational Approaches to Arabic-English Code-Switching

Natural Language Processing (NLP) is a vital computational method for addressing language processing

ERROR ANALYSIS OF TIGRINYA– ENGLISH MACHINE  TRANSLATION SYSTEMS

ERROR ANALYSIS OF TIGRINYA– ENGLISH MACHINE  TRANSLATION SYSTEMS

Poster presented at the Deep Learning Indaba 2023 by Negasi Abadi

ChatGPT for Arabic-English Translation: Evaluating the Accuracy ChatGPT للترجمة من العربية إلى الإنجليزية: تقييم الدقة ChatGPT pour la traduction arabe-anglais : évaluation de l'exactitude ChatGPT para traducción árabe-inglés: evaluación de la precisión

Abstract Cross-cultural communication has become more accessible with the advent of ChatGPT as a tra