This project contains the complete data, code, and supplementary materials underlying the associated preprint. All primary diagnostic judgments, the full coded dataset, analysis scripts, and every statistical figure and table reported in the paper are included so that all quantitative claims can be independently verified and reproduced.
Contents:
/Supplementary_Materials/ — full appendices (A–G): methods, language sample composition, inter-rater reliability, speaker demographics, per-language syntactic/morphological diagnostic results, and the complete quantitative validation methodology (PCA, MCA, jackknife, bootstrap, and independent Universal Dependencies corpus validation).
/Primary_Data/ — raw, per-speaker diagnostic judgments and morphological corroboration data for each of the six Core languages (Balti, English, Hindi-Urdu, Italian, Swahili, Warlpiri), traceable to every quantitative claim in the manuscript.
/Final_Dataset/ — the complete coded 82-language dataset used in all statistical analyses.
/Statistical_Analysis_Outputs/ — all figures and tables reported in the manuscript and supplementary materials, generated directly from the scripts below.
/Scripts/ — the R script producing all statistical analyses and figures (PCA, MCA, jackknife, bootstrap, and diagnostic figure generation) from the dataset in /Final_Dataset/, and a separate Python script implementing the independent Universal Dependencies φ-state validation pipeline (Appendix G).