This dataset contains sample outputs and evaluation scores from the study “A Test of Meaning, Form, and Culture in Kurmanji: An Evaluation of Large Language Models’ Performance.”
It includes responses generated by four large language models—ChatGPT-4o, Gemini, Claude, and Grok—in reaction to prompts written in Kurmanji, a morphologically rich and culturally embedded Kurdish language.
Each model output was scored by human annotators based on grammatical correctness, semantic accuracy, and cultural appropriateness. The dataset also includes evaluation notes for qualitative insight.
This sample represents a subset of the full analysis described in the manuscript.
The dataset is shared under the CC-BY 4.0 license, and may be reused for research, teaching, or further model evaluation in low-resource language contexts.