Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Ranks without resolution: data and code for a measurement audit of a multilingual LLM benchmark

Domaine:

natural language processing

Type de record:

dataset
Créateur:
Kan
Éditeur:
Zenodo
Hôte:avatar
Data, code and results for a measurement audit of HELM's African-language MMLU and Winogrande suite, accompanying the preprint 'Ranks without resolution: how much of a multilingual benchmark's language ordering is estimable?'. The deposit contains the retrieval scripts, the analysis pipeline, the item response theory implementation with its golden self-test, the permutation-null and linking code, the three figures, and every results table reported in the paper. All input data are public and no credentials are needed at any stage. Response matrices were retrieved from the Stanford CRFM HELM public results bucket. The deposit holds item identifiers, scored binary outcomes and derived statistics only, and reproduces every number reported in the paper. Upstream scenarios keep their own licences (MMLU MIT, Winogrande Apache-2.0), the African-language translations derive from the release of Alhanai et al. (2025), and no item text is reproduced or redistributed here.

Visit

doi.org

Tags

benchmark validitydifferential item functioningmultilingual evaluationmeasurement precisionlarge language modelsitem response theory

Licenses

Creative Commons Attribution 4.0 Internationalhttps://creativecommons.org/licenses/by/4.0/legalcode