Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

MoVoC: Morphology-Aware Subword Construction for Geez Script Languages

Domain:

natural language processing

Record type:

paperdatasetsoftware
Creator:
TekFazNej
Host:avatar
Subword-based tokenization methods often fail to preserve morphological boundaries, a limitation especially pronounced in low-resource, morphologically complex languages such as those written in the Geez script. To address this, we present MoVoC (Morpheme-aware Subword Vocabulary Construction) and train MoVoC-Tok, a tokenizer that integrates supervised morphological analysis into the subword vocabulary. This hybrid segmentation approach combines morpheme-based and Byte Pair Encoding (BPE) tokens to preserve morphological integrity while maintaining lexical meaning. To tackle resource scarcity, we curate and release manually annotated morpheme data for four Geez script languages and a morpheme-aware vocabulary for two of them. While the proposed tokenization method does not lead to significant gains in automatic translation quality, we observe consistent improvements in intrinsic metrics, MorphoScore, and Boundary Precision, highlighting the value of morphology-aware segmentation in enhancing linguistic fidelity and token efficiency. Our morpheme-annotated datasets and tokenizer will be publicly available to support further research in low-resource, morphologically rich languages. Our code and data are available on GitHub: github.com This submission is approximately 10 pages in length and includes 1 figure and 6 tables

Visit

arxiv.org

Tasks

machine translation

Tags

Computation and LanguageArtificial IntelligenceI.2.7; I.2.6; H.3.3

Similar

Morphology-aware Subword Segmentation for Zero-shot Cross-lingual Transfer in Low-resource LanguagesMorphology-Aware Multi-Granularity Representation Learning for Agglutinative LanguagesHierarchical Multi Task Learning with Subword Contextual Embeddings for Languages with Rich MorphologyTabuLM: Morphology-Aware Tabular Pre-training for Low-Resource LanguagesAncient Geez script recognition using deep learningSubword Segmental Language Modelling for Nguni Languages

Morphology-aware Subword Segmentation for Zero-shot Cross-lingual Transfer in Low-resource Languages

Multilingual modelling can improve machine translation for low-resource languages, partly through sh

Morphology-Aware Multi-Granularity Representation Learning for Agglutinative Languages

Low-resource agglutinative languages, characterized by rich morphological inflection and severe voca

Hierarchical Multi Task Learning with Subword Contextual Embeddings for Languages with Rich Morphology

Morphological information is important for many sequence labeling tasks in Natural Language Processi

TabuLM: Morphology-Aware Tabular Pre-training for Low-Resource Languages

We present TabuLM, the first language model pre-trained on Kinyarwanda tabular data. Kinyarwanda is

Ancient Geez script recognition using deep learning

Subword Segmental Language Modelling for Nguni Languages

Subword Segmental Language Modelling for Nguni Languages

Poster presented at the Deep Learning Indaba 2022 by Jan Buys