Data, code and results for a measurement audit of HELM's African-language MMLU and Winogrande suite, accompanying the preprint 'Ranks without resolution: how much of a multilingual benchmark's language ordering is estimable?'. The deposit contains the retrieval scripts, the analysis pipeline, the item response theory implementation with its golden self-test, the permutation-null and linking code, the three figures, and every results table reported in the paper.
All input data are public and no credentials are needed at any stage. Response matrices were retrieved from the Stanford CRFM HELM public results bucket. The deposit holds item identifiers, scored binary outcomes and derived statistics only, and reproduces every number reported in the paper. Upstream scenarios keep their own licences (MMLU MIT, Winogrande Apache-2.0), the African-language translations derive from the release of Alhanai et al. (2025), and no item text is reproduced or redistributed here.