
Evaluation materials accompanying the paper "What Single-Reference BLEU Cannot Show: A Pre-Registered Blind Human Evaluation of Uzbek–English Machine Translation". Contains the frozen 72-sentence corpus (three stylistic layers), 216 system outputs, the blinding key, anonymized human ratings from two independent raters, per-item BLEU/chrF scores, the frozen primary analysis script, and the dated preregistration plan with amendments.