Measurement dataset for The African Language Tax — the first systematic audit of the subword tokenization penalty imposed on African languages by frontier large language models.
Every row is one (language, tokenizer, corpus) triple, with fertility, English-relative premium, and confidence intervals computed from a parallel corpus using sum-then-divide aggregation.
Property
Value