Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

The analysis of customized tokenizer for the development of isiNdebele language model

Creator:
ProThi
Publisher:
IEEE
Host:

Visit

doi.org

Tasks

language modeling

Languages

NdebeleNdebele

Licenses

https://doi.org/10.15223/policy-029https://doi.org/10.15223/policy-037

Similar

NCHLT isiNdebele RoBERTa language modelIsiNdebele monolingual language model using custom NguniTokenizerPOLYLAB: A CUSTOMIZED MULTI-LANGUAGE DEVELOPMENT ENVIRONMENT FOR COMPUTER SCIENCE STUDENTS IN NIGERIAN POLYTECHNICSMehranLM-Tokenizer: A Natively-Trained Tokenizer for Sindhi Language Modellinglabrijisaad/Sentiment-Analysis-model-for-the-Wolof-languageThe Efficiency of IsiNdebele Part of Speech Tagger: A Quantitative Analysis

NCHLT isiNdebele RoBERTa language model

Contextual masked language model based on the RoBERTa architecture (Liu et al., 2019). The model is

IsiNdebele monolingual language model using custom NguniTokenizer

POLYLAB: A CUSTOMIZED MULTI-LANGUAGE DEVELOPMENT ENVIRONMENT FOR COMPUTER SCIENCE STUDENTS IN NIGERIAN POLYTECHNICS

In this paper, a customized multi-language development environment (MLDE), PolyLAB, is developed for

MehranLM-Tokenizer: A Natively-Trained Tokenizer for Sindhi Language Modelling

Sindhi, an Indo-Aryan language spoken by tens of millions of people, remains severely underrepresent

labrijisaad/Sentiment-Analysis-model-for-the-Wolof-language

In this notebook, I tried to create a sentiment analysis model for Wolof language. # 📈 `Sentiment

The Efficiency of IsiNdebele Part of Speech Tagger: A Quantitative Analysis