Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

miloudbelarebia/arabic-tax

Domaine:

natural language processing
Créateur:
mil
Hôte:
Arabic Costs Double on AI — why every Arabic-script language (Darija, MSA, Persian, Urdu) pays a 2x token tax on ChatGPT, Claude & Gemini. # arabic-tax > **The same sentence costs 2× more tokens in Arabic than in English on AI.** > Not an opinion. The mechanical result of two design choices made in 1992 and 2019. > Affects every Arabic-script language — Darija, Algerian, Egyptian, MSA, Persian, Urdu — across **ChatGPT, Claude, Gemini**, and every major LLM. --- ## TL;DR I downloaded OpenAI's official tokenizer file (`o200k_base`, 200,019 entries, used by GPT-4o / GPT-4.1 / GPT-5.x / o1 / o3 — verified on platform.openai.com, all return identical token counts). I measured exactly how it splits the vocabulary by writing system. **Latin gets 67.4% of the dictionary. Arabic gets 4.0%. The ratio is 17 to 1.** Direct consequence: the same idea takes about twice as many tokens in Arabic as in English on ChatGPT — more API cost, smaller usable context window, slower answers, worse output quality. This repo holds the **reproducible Python notebook**, the **infographic**, and the **historical analysis** of why this happens. --- ## Why does this happen? A 32-year-old design story This isn't a bug. It's the **stacked consequence of two design choices** — one from 1992, one from 2019. Both were defensible at the time. Neither was malicious. But together, they produce today's tax on Arabic-script users. ### 🧱 Layer 1 — UTF-8 (1992): the foundation **What it is.** UTF-8 is the standard that decides how text is stored as bytes on disk. Every text file, every webpage, every API call uses it. It was designed by Ken Thompson and Rob Pike in September 1992, on a placemat in a New Jersey diner. **The trade-off they made.** UTF-8 had to stay 100% backward-compatible with ASCII (the old 128-character English-only standard from the 1960s). So the design was: the 128 ASCII characters keep their old 1-byte encoding. Everything else needs more bytes. **The mechanical result, as written in RFC 3629:** | Unicode range | Byte cost | Examples | |---|---:|---| | U+0000 – U+007F | **1 byte** | `A`, `z`, digit …

Visit

github.com

Languages

Arabic, Algerian Spoken

Tags

ai-fairnessalgerian-arabicarabicarabic-nlpchatgptdarijadata-journalismegyptian-arabicjupyter-notebooklinguistic-diversity+10

Similaires

Liberia-Tax/Liberia-Tax-MicrosimulationTax Education to Improve Tax Knowledge and Tax Morale of Young Adults in Cameroon.TAX AUDIT PRACTICES AND TAX COMPLIANCE IN NIGERIATax Policy in WAEMU: Tax Coordination or Competition?DIGITAL TAX ADMINISTRATION AND TAX COMPLIANCE BEHAVIOUR: AN EVALUATION OF THE NIGERIA TAX ACT 2025Are Tax Penalties Effective Enough in Combating Tax Evasion?

Liberia-Tax/Liberia-Tax-Microsimulation

Demonstration repository for Botswana microsimulation model =======================================

Tax Education to Improve Tax Knowledge and Tax Morale of Young Adults in Cameroon.

In this research, we collect follow-up data for the randomized survey experiments pre-registered at

TAX AUDIT PRACTICES AND TAX COMPLIANCE IN NIGERIA

This study examined the effect of tax audit practices on tax compliance in Nigeria, with emphasis on

Tax Policy in WAEMU: Tax Coordination or Competition?

The objective of this paper is to examine the effect of digitalisation on tax revenue mobilisation i

DIGITAL TAX ADMINISTRATION AND TAX COMPLIANCE BEHAVIOUR: AN EVALUATION OF THE NIGERIA TAX ACT 2025

The study was motivated by the enactment of the Nigeria Tax Act (NTA) 2025, which among other things

Are Tax Penalties Effective Enough in Combating Tax Evasion?

This paper examined whether Tax Penalties is Effective in Combating Tax Evasion in Nigeria. The stud