A Python script demonstrating the 'Tokenization Tax' by comparing GPT-4's token efficiency across English, Swahili, and Sheng
# Semantic Dilution: Tokenization Tax in Swahili and Sheng
### The Problem: Semantic Dilution
Modern AI tokenizers are trained predominantly on English text. Because of this, they efficiently recognize whole English words (keeping the Tokens-per-Word ratio close to 1.0).
However, when processing African languages like Swahili or dynamic slang like Sheng, the tokenizer applies English character-pairing rules to foreign structures. This forces the machine to shatter these words into tiny, unrecognizable fragments. This fragmentation heavily inflates the Tokens-per-Word ratio, leading to a "Tokenization Tax" where Swahili and Sheng cost more compute power and API budget to process than English.
### The Methodology
To analyze the token fragmentation, this project uses:
* **Library:** OpenAI's `tiktoken` to process the text.
* **Tokenizer:** The `cl100k_base` model (the same one used by GPT-4).
* **Metric:** The script calculates a 'Tokens-per-Word Ratio' by dividing the total tokens generated by the original word count.
* **Visualization:** ANSI escape codes alternate terminal background colors to make the exact sub-word splits visually obvious.
### The Results
Running the exact same retail description through the `cl100k_base` tokenizer reveals a massive difference in efficiency. Here is the raw terminal output:
```text
--- English ---
Words: 6 | Tokens: 6 | Ratio: 1.00
Fragmentation: High quality leather shoes for men
--- Swahili ---
Words: 9 | Tokens: 20 | Ratio: 2.22
Fragmentation: Vi atu vya ngozi vya ubora wa juu kwa wanaume
--- Sheng ---
Words: 7 | Tokens: 12 | Ratio: 1.71
Fragmentation: F egi za ng ware high quality za mab oys