A 150-item demonstration corpus of native-written Arabic across three varieties — Modern Standard Arabic, Palestinian Levantine and Egyptian (50 items each) — each item carrying an English gloss. Every sentence was written from scratch by a first-language speaker of the target variety. Nothing in this corpus was scraped from the web, machine-translated from English, or generated by a language model.
The deposit's purpose is not the data alone. Most Arabic training material is adapted from English by translation, and dialectal coverage in particular leans on machine translation; the resulting text is grammatical but reads as no Arabic speaker writes. A clean dataset, however, cannot prove its own origin — a language model also produces clean output. What distinguishes human authorship is the record of a human writer making human mistakes and a human reviewer catching them, and that record is normally discarded.
This deposit preserves it. The accompanying provenance record (PROVENANCE.md) documents the contributor brief and its explicit prohibition on copying, translating and AI generation; the review passes and the objections raised; the exact cells that changed between submission and release; and a measured diversity target that was deliberately left unmet because the native speaker judged the shorter wording unnatural and her judgement was kept over the metric. It also records what the chain does not cover: the Palestinian Levantine set carries no before/after correction artifact, and two published figures rest on reviewer judgement rather than a mechanical measure.
All published figures were independently recomputed from the released CSV files on 15 July 2026 and are documented in section 10 of the provenance record. SHA-256 checksums cover both the deposited files and, as a public hash commitment, the unpublished internal audit trail.
This is a demonstration sample, not a training-scale corpus: 50 items per variety, two authors. It exists so that quality can be judged before anything is commissioned.
The same three sets are also published individually on Hugging Face (access-gated) and Kaggle (open download).