Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

AraLingBench A Human-Annotated Benchmark for Evaluating Arabic Linguistic Capabilities of Large Language Models

Domain:

natural language processing

Record type:

datasetpaper

We present AraLingBench: a fully human annotated benchmark for evaluating the Arabic linguistic competence of large language models (LLMs). The benchmark spans five core categories: grammar, morphology, spelling, reading comprehension, and syntax, through 150 expert-designed multiple choice questions that directly assess structural language understanding. Evaluating 35 Arabic and bilingual LLMs reveals that current models demonstrate strong surface level proficiency but struggle with deeper grammatical and syntactic reasoning. AraLingBench highlights a persistent gap between high scores on knowledge-based benchmarks and true linguistic mastery, showing that many models succeed through memorization or pattern recognition rather than authentic comprehension. By isolating and measuring fundamental linguistic skills, AraLingBench provides a diagnostic framework for developing Arabic LLMs. The full evaluation code is publicly available on GitHub.

Visit

arxiv.orghuggingface.cogithub.comdataset on huggingface

Tags

evaluationbenchmarkmultiple-choice

Similar

UrBLiMP: A Benchmark for Evaluating the Linguistic Competence of Large Language Models in UrduEvaluating the Capabilities of Large Language Models for Multi-label Emotion UnderstandingBeyond Fluency: A Rubric-Based Benchmark for Evaluating Saudi Dialect and Cultural Competence in Large Language ModelsM3Exam: A Multilingual, Multimodal, Multilevel Benchmark for Examining Large Language ModelsBALSAM: A Platform for Benchmarking Arabic Large Language ModelsBenchmark Dataset for DiaLex, A Benchmark for Evaluating Multidialectal Arabic Word Embeddings

UrBLiMP: A Benchmark for Evaluating the Linguistic Competence of Large Language Models in Urdu

Multilingual Large Language Models (LLMs) have shown remarkable performance across various languages

Evaluating the Capabilities of Large Language Models for Multi-label Emotion Understanding

Large Language Models (LLMs) show promising learning and reasoning abilities. Compared to other NLP

Beyond Fluency: A Rubric-Based Benchmark for Evaluating Saudi Dialect and Cultural Competence in Large Language Models

Large language models are increasingly deployed in Arabic-speaking markets, yet standard benchmarks

M3Exam: A Multilingual, Multimodal, Multilevel Benchmark for Examining Large Language Models

Despite the existence of various benchmarks for evaluating natural language processing models, we ar

BALSAM: A Platform for Benchmarking Arabic Large Language Models

The impressive advancement of Large Language Models (LLMs) in English has not been matched across al

Benchmark Dataset for DiaLex, A Benchmark for Evaluating Multidialectal Arabic Word Embeddings

Word embeddings are a core component of modern natural language processing systems, making the ability to thoroughly evaluate them a vital task. We describe DiaLex, a benchmark for intrinsic evaluation of dialectal Arabic word embeddings. DiaLex covers five importa