Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Ranks without resolution: data and code for a measurement audit of a multilingual LLM benchmark

Domain:

natural language processing

Record type:

dataset
Creator:
Kan
Publisher:
Zenodo
Host:avatar
Data, code and results for a measurement audit of HELM's African-language MMLU and Winogrande suite, accompanying the preprint 'Ranks without resolution: how much of a multilingual benchmark's language ordering is estimable?'. The deposit contains the retrieval scripts, the analysis pipeline, the item response theory implementation with its golden self-test, the permutation-null and linking code, the three figures, and every results table reported in the paper. All input data are public and no credentials are needed at any stage. Response matrices were retrieved from the Stanford CRFM HELM public results bucket. The deposit holds item identifiers, scored binary outcomes and derived statistics only, and reproduces every number reported in the paper. Upstream scenarios keep their own licences (MMLU MIT, Winogrande Apache-2.0), the African-language translations derive from the release of Alhanai et al. (2025), and no item text is reproduced or redistributed here.

Visit

doi.org

Tags

benchmark validitydifferential item functioningmultilingual evaluationmeasurement precisionlarge language modelsitem response theory

Licenses

Creative Commons Attribution 4.0 Internationalhttps://creativecommons.org/licenses/by/4.0/legalcode

Similar

Minionese: Comprehensive Benchmark and Mechanistic Study of Multilingual LLM SafetyData and Code for: Improving Willingness-to-Pay Elicitation by Including a Benchmark GoodMultiGlobeQA: A Multilingual and Globally Diverse Benchmark for Geospatial ReasoningLinCE: A Centralized Benchmark for Linguistic Code-switching EvaluationPOLAR: A Benchmark for Multilingual, Multicultural, and Multi-Event Online PolarizationValidating LLM Annotation Without a Gold Standard: A Three-Way Design Applied to Hybrid Arabic–French Corpora — annotation labels, code and materials

Minionese: Comprehensive Benchmark and Mechanistic Study of Multilingual LLM Safety

Safety alignment in large language models remains brittle across languages: prompts reliably refused

Data and Code for: Improving Willingness-to-Pay Elicitation by Including a Benchmark Good

Abstract: We propose and validate a simple way to augment the standard Becker-DeGroot-Marscha

MultiGlobeQA: A Multilingual and Globally Diverse Benchmark for Geospatial Reasoning

Geospatial reasoning, i.e., computing distances, containment, and other spatial relations over real-

LinCE: A Centralized Benchmark for Linguistic Code-switching Evaluation

Recent trends in NLP research have raised an interest in linguistic code-switching (CS); modern appr

POLAR: A Benchmark for Multilingual, Multicultural, and Multi-Event Online Polarization

Online polarization poses a growing challenge for democratic discourse, yet most computational socia

Validating LLM Annotation Without a Gold Standard: A Three-Way Design Applied to Hybrid Arabic–French Corpora — annotation labels, code and materials

Annotation labels, analysis code and coding materials for a three-way validation study of LLM-assist