Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

The Qiyas Benchmark: Measuring ChatGPT Mathematical and Language Understanding in Arabic

Domaine:

natural language processing

Type de record:

paperdataset
Créateur:
Al-Al-
Hôte:avatar
Despite the growing importance of Arabic as a global language, there is a notable lack of language models pre-trained exclusively on Arabic data. This shortage has led to limited benchmarks available for assessing language model performance in Arabic. To address this gap, we introduce two novel benchmarks designed to evaluate models' mathematical reasoning and language understanding abilities in Arabic. These benchmarks are derived from a General Aptitude Test (GAT) called Qiyas exam, a standardized test widely used for university admissions in Saudi Arabia. For validation purposes, we assess the performance of ChatGPT-3.5-trubo and ChatGPT-4 on our benchmarks. Our findings reveal that these benchmarks pose a significant challenge, with ChatGPT-4 achieving an overall average accuracy of 64%, while ChatGPT-3.5-trubo achieved an overall accuracy of 49% across the various question types in the Qiyas benchmark. We believe the release of these benchmarks will pave the way for enhancing the mathematical reasoning and language understanding capabilities of future models tailored for the low-resource Arabic language.

Visit

arxiv.org

Tags

Computation and LanguageArtificial Intelligence

Similaires

TUMLU: A Unified and Native Language Understanding Benchmark for Turkic LanguagesNSMQ Riddles: A Benchmark of Scientific and Mathematical Riddles for Quizzing Large Language ModelsArabicMMLU: Assessing Massive Multitask Language Understanding in ArabicJEEM: Vision-Language Understanding in Four Arabic DialectsFleurs-SLU: A Massively Multilingual Benchmark for Spoken Language UnderstandingTARIC-SLU: A Tunisian Benchmark Dataset for Spoken Language Understanding

TUMLU: A Unified and Native Language Understanding Benchmark for Turkic Languages

Being able to thoroughly assess massive multi-task language understanding (MMLU) capabilities is ess

NSMQ Riddles: A Benchmark of Scientific and Mathematical Riddles for Quizzing Large Language Models

Large Language Models (LLMs) have shown good performance on various science educational benchmarks,

ArabicMMLU: Assessing Massive Multitask Language Understanding in Arabic

The focus of language model evaluation has transitioned towards reasoning and knowledge-intensive ta

JEEM: Vision-Language Understanding in Four Arabic Dialects

We introduce JEEM, a benchmark designed to evaluate Vision-Language Models (VLMs) on visual understa

Fleurs-SLU: A Massively Multilingual Benchmark for Spoken Language Understanding

Spoken language understanding (SLU) is indispensable for half of all living languages that lack a fo

TARIC-SLU: A Tunisian Benchmark Dataset for Spoken Language Understanding