Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Luganda Tokenizer Evaluation

Domain:

natural language processing

Record type:

dataset
Creator:
Cra
Host:
Cross-lingual evaluation of 9 decoder-only LLMs on English (CoNLL-2003, 5,396 sentences) and Luganda (MasakhaNER, 6,055 sentences), quantifying the "tokenizer tax" that low-resource languages pay when using tokenizers designed for English.

Visit

huggingface.co

Languages

Ganda

Tags

lugandatokenizerevaluationlow-resourceafrican-languagesmorphologybpcfertilitycross-lingual

Licenses

apache-2.0

Similar

Kibalama/luganda-tokenizerrashid0784/luganda-dialect-tokenizermewaeltsegay/tokenizerMehranLM-Tokenizer: A Natively-Trained Tokenizer for Sindhi Language Modelling

Kibalama/luganda-tokenizer

rashid0784/luganda-dialect-tokenizer

mewaeltsegay/tokenizer

Tigrinya Language Tokenizers # Tigrinya Multi-Type Tokenizers for LLM Training A comprehensive col

MehranLM-Tokenizer: A Natively-Trained Tokenizer for Sindhi Language Modelling

Sindhi, an Indo-Aryan language spoken by tens of millions of people, remains severely underrepresent