Morphological analysis constitutes a foundational task in natural language processing (NLP). Sub-word algorithms have demonstrated efficacy across various NLP tasks. particularly for morphologically rich languages. However, their application to tonal languages with non-concatenative morphology remains underexplored in literature. This study investigates the application of sub-word tokenization algorithms for morphological analysis of Yoruba, a tonal Niger-Congo language with rich agglutinative morphology. The experiment and evaluation were based on three prominent sub-word algorithms which are Byte-Pair Encoding (BPE), Byte-level BPE and WordPiece. These were used to generate morphological wordforms from constituent morphemes. The evaluation was conducted on a curated Yoruba morphological dataset. The measure of performance of each algorithm is based on segmentation accuracy and computational efficiency. The results show that WordPiece splits words accurately (74.8%), but it uses more computing power and takes longer to run (28.7 seconds training time, 6,340 words/sec). Byte-level BPE strikes a better balance between speed and accuracy, especially when handling Yoruba’s accent marks. These results help developers choose the right tokenizer for Yoruba language projects.
Keywords- Morphological Analysis, Yoruba Text, Sub-Word Algorithms, Byte-Pair Encoding, Tonal
Aliyu, E.O. (2026): Morphological Analysis of Yoruba Text Using Sub-Word Algorithms. Journal of Advances in Mathematical & Computational Science. Vol. 14, No. 2. Pp 1-8 Available online at
isteams.net.
dx.doi.org