Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

An Empirical Study of Preprocessing and Vocabulary Effects in Myanmar Unigram Tokenization

Domain:

natural language processing

Record type:

paper
Creator:
Kha
Publisher:
Zenodo
Host:avatar
This paper presents an empirical study of preprocessing strategies and vocabulary sizes for Myanmar Unigram tokenization. The study evaluates multiple preprocessing variants, including Raw, No-space, Half-mixed, syllable-based segmentation, and ZWSP-based segmentation under different vocabulary capacities. The experiments analyze compression efficiency, fragmentation behavior, fertility, token compactness, and vocabulary utilization. Results show that larger vocabularies generally reduce fragmentation and improve token compactness, while preprocessing strategies strongly influence linguistic stability. In particular, ZWSP-based segmentation demonstrated lower fragmentation behavior while preserving stable token boundaries. The study further shows that smaller vocabularies may occasionally produce linguistically incomplete Myanmar fragments and isolated combining marks despite stable decoding integrity. This work provides one of the first fragmentation-aware empirical analyses of Myanmar Unigram tokenization behavior across multiple preprocessing conditions.

Visit

doi.orgzenodo.org

Tasks

text normalization

Tags

Myanmar NLPMyanmar TokenizationSentencePieceUnigram TokenizationLow-resource NLPSubword TokenizationMyanmar LanguageToken FragmentationZWSPNatural Language Processing

Licenses

Creative Commons Attribution 4.0 Internationalhttps://creativecommons.org/licenses/by/4.0/legalcode© 2026, Khant Sint Heinnhttp://rightsstatements.org/vocab/InC/1.0/

Similar

Trans-Tokenization and Cross-lingual Vocabulary Transfers: Language Adaptation of LLMs for Low-Resource NLPBilingual Proficiency Effects in Paired Associate Learning of Vocabulary in an Unfamiliar LanguageAn Exploration of Vocabulary Size and Transfer Effects in Multilingual Language Models for African LanguagesMariiaChugaeva/NLP-Case-Study-Multilingual-TokenizationAn Empirical Study of University Education and Graduate Employability in TanzaniaVelkamez/tamazight-unigram

Trans-Tokenization and Cross-lingual Vocabulary Transfers: Language Adaptation of LLMs for Low-Resource NLP

The development of monolingual language models for low and mid-resource languages continues to be hi

Bilingual Proficiency Effects in Paired Associate Learning of Vocabulary in an Unfamiliar Language

This data set is reported in an article with the same title to be published in Bilingualism: Languag

An Exploration of Vocabulary Size and Transfer Effects in Multilingual Language Models for African Languages

Multilingual pretrained language models have been shown to work well on many languages, even those they were not originally pretrained on. Despite their empirical success in downstream tasks, there is still a gap in understanding of "what makes them tick''. In this

MariiaChugaeva/NLP-Case-Study-Multilingual-Tokenization

This project investigates how tokenization choices affect multilingual NLP under a low-resource setu

An Empirical Study of University Education and Graduate Employability in Tanzania

Few would challenge the proposition that employability skills are fundamental for success in the wor

Velkamez/tamazight-unigram