Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

KurdishMCQ: A multiple-choice question dataset for the Kurdish language (Sorani dialect)

Domain:

natural language processing

Record type:

dataset
Creator:
AhmRas
Editor:
Sul
Publisher:
Men
Host:avatar
This dataset is the first large-scale, publicly available collection of multiple-choice questions in the Kurdish language (Sorani dialect), helping fill an important gap in evaluating large language models (LLMs) for low-resource languages. It includes 17,233 questions gathered from Kurdish middle and high school materials, such as textbooks and official exams, all written in Sorani using the Perso-Arabic script. The dataset was carefully created and reviewed by native Kurdish speakers to ensure accuracy and consistency. Each entry follows a clear and structured format, including fields such as the question text, four answer options (A–D), the correct answer label, and metadata like subject, grade level, and group (e.g., STEM). The dataset covers a wide range of subjects, with the largest portions coming from Biology (2636 questions), Chemistry (2326), Kurdish grammar (1591), Natural Science (1537), and Kurdish literature (1488), along with other areas like Physics, Math, History, and Economics. In terms of grade distribution, most questions are from grade 12 (9453), followed by grades 9, 10, and 11, with a smaller portion categorized as “other”. This resource can be used for tasks like question answering, knowledge evaluation, and reasoning, and is well-suited for training, fine-tuning, and benchmarking language models, especially in multilingual and low-resource settings.

Visit

doi.org

Tasks

question answering

Tags

Natural Language ProcessingLarge Language Model

Licenses

info:eu-repo/semantics/openAccessCreative Commons Attribution 4.0 Internationalhttps://creativecommons.org/licenses/by/4.0/legalcode

Similar

Named Entity Recognition for the Kurdish Sorani Language: Dataset Creation and Comparative AnalysisKHLD: A Large-Scale Benchmark of the Kurdish Handwritten Lines Dataset for Low-Resource Central Kurdish (Sorani)Kurdish (Sorani) Speech to Text: Presenting an Experimental DatasetAutomatic Text Summarization (ATS) for Research Documents in Sorani KurdishA paper-based cheat-resistant multiple-choice question system with automated gradingUsing Punkt for Sentence Segmentation in non-Latin Scripts: Experiments on Kurdish (Sorani) Texts

Named Entity Recognition for the Kurdish Sorani Language: Dataset Creation and Comparative Analysis

This work contributes towards balancing the inclusivity and global applicability of natural language

KHLD: A Large-Scale Benchmark of the Kurdish Handwritten Lines Dataset for Low-Resource Central Kurdish (Sorani)

The Kurdish Handwritten Lines Dataset (KHLD) is a large-scale image dataset aiming to facilitate han

Kurdish (Sorani) Speech to Text: Presenting an Experimental Dataset

We present an experimental dataset, Basic Dataset for Sorani Kurdish Automatic Speech Recognition (B

Automatic Text Summarization (ATS) for Research Documents in Sorani Kurdish

Extracting concise information from scientific documents aids learners, researchers, and practitione

A paper-based cheat-resistant multiple-choice question system with automated grading

This paper focuses on how to reduce cheating and minimize errors while automatically grading paper-b

Using Punkt for Sentence Segmentation in non-Latin Scripts: Experiments on Kurdish (Sorani) Texts

Segmentation is a fundamental step for most Natural Language Processing tasks. The Kurdish language