Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Tokenization Disparities as Infrastructure Bias: How Subword Systems Create Inequities in LLM Access and Efficiency

Domaine:

natural language processing

Type de record:

paper
Créateur:
TekNej
Hôte:avatar
Tokenization disparities pose a significant barrier to achieving equitable access to artificial intelligence across linguistically diverse populations. This study conducts a large-scale cross-linguistic evaluation of tokenization efficiency in over 200 languages to systematically quantify computational inequities in large language models (LLMs). Using a standardized experimental framework, we applied consistent preprocessing and normalization protocols, followed by uniform tokenization through the tiktoken library across all language samples. Comprehensive tokenization statistics were collected using established evaluation metrics, including Tokens Per Sentence (TPS) and Relative Tokenization Cost (RTC), benchmarked against English baselines. Our cross-linguistic analysis reveals substantial and systematic disparities: Latin-script languages consistently exhibit higher tokenization efficiency, while non-Latin and morphologically complex languages incur significantly greater token inflation, often 3-5 times higher RTC ratios. These inefficiencies translate into increased computational costs and reduced effective context utilization for underrepresented languages. Overall, the findings highlight structural inequities in current AI systems, where speakers of low-resource and non-Latin languages face disproportionate computational disadvantages. Future research should prioritize the development of linguistically informed tokenization strategies and adaptive vocabulary construction methods that incorporate typological diversity, ensuring more inclusive and computationally equitable multilingual AI systems. 6 pages 4 figures

Visit

arxiv.org

Tags

Computation and LanguageArtificial IntelligenceI.2.7; I.2.1; H.3.3; F.2.2

Similaires

Addressing Voids: How Digital Start-ups in Kenya Create Market InfrastructureThe Token Tax: Systematic Bias in Multilingual TokenizationSubword Tokenization Strategies in Cross-Lingual Retrieval for Kinyarwanda: mBERT Versus Monolingual RoBERTa on XTREMEGlobal Disparities in Pediatric Radiotherapy Infrastructure and Access: A Scoping Review With a Focused Analysis of MoroccoFairness or Fluency? An Investigation into Language Bias of Pairwise LLM-as-a-JudgeHow to create and use your ORCID account?

Addressing Voids: How Digital Start-ups in Kenya Create Market Infrastructure

Abstract The chapter introduces market-enabling digital platforms as a new concept. By using a mark

The Token Tax: Systematic Bias in Multilingual Tokenization

Tokenization inefficiency imposes structural disadvantages on morphologically complex, low-resource

Subword Tokenization Strategies in Cross-Lingual Retrieval for Kinyarwanda: mBERT Versus Monolingual RoBERTa on XTREME

Pre-trained multilingual language models (e.g., mBERT, XLM-RoBERTa) have significantly advanced the

Global Disparities in Pediatric Radiotherapy Infrastructure and Access: A Scoping Review With a Focused Analysis of Morocco

ABSTRACT Radiotherapy (RT) is a crucial component of childhood cancer management

Fairness or Fluency? An Investigation into Language Bias of Pairwise LLM-as-a-Judge

Recent advances in Large Language Models (LLMs) have incentivized the development of LLM-as-a-judge,

How to create and use your ORCID account?

This tutorial is created as part of the INASP's Project: Digital Hub in East Africa. &