Logo Lanfrica

Oak-ke/Semantic-Dilution

Domain:

natural language processing

Record type:

software
Creator:
Oak
Host:
A Python script demonstrating the 'Tokenization Tax' by comparing GPT-4's token efficiency across English, Swahili, and Sheng # Semantic Dilution: Tokenization Tax in Swahili and Sheng ### The Problem: Semantic Dilution Modern AI tokenizers are trained predominantly on English text. Because of this, they efficiently recognize whole English words (keeping the Tokens-per-Word ratio close to 1.0). However, when processing African languages like Swahili or dynamic slang like Sheng, the tokenizer applies English character-pairing rules to foreign structures. This forces the machine to shatter these words into tiny, unrecognizable fragments. This fragmentation heavily inflates the Tokens-per-Word ratio, leading to a "Tokenization Tax" where Swahili and Sheng cost more compute power and API budget to process than English. ### The Methodology To analyze the token fragmentation, this project uses: * **Library:** OpenAI's `tiktoken` to process the text. * **Tokenizer:** The `cl100k_base` model (the same one used by GPT-4). * **Metric:** The script calculates a 'Tokens-per-Word Ratio' by dividing the total tokens generated by the original word count. * **Visualization:** ANSI escape codes alternate terminal background colors to make the exact sub-word splits visually obvious. ### The Results Running the exact same retail description through the `cl100k_base` tokenizer reveals a massive difference in efficiency. Here is the raw terminal output: ```text --- English --- Words: 6 | Tokens: 6 | Ratio: 1.00 Fragmentation: High quality leather shoes for men --- Swahili --- Words: 9 | Tokens: 20 | Ratio: 2.22 Fragmentation: Vi atu vya ngozi vya ubora wa juu kwa wanaume --- Sheng --- Words: 7 | Tokens: 12 | Ratio: 1.71 Fragmentation: F egi za ng ware high quality za mab oys