Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Enhancing African low-resource languages: Swahili data for language modelling

Domaine:

natural language processing

Type de record:

paper
Language modelling using neural networks requires adequate data to guarantee quality word representation which is important for natural language processing (NLP) tasks. However, African languages, Swahili in particular, have been disadvantaged and most of them are classified as low resource languages because of inadequate data for NLP. In this article, we derive and contribute unannotated Swahili dataset, Swahili syllabic alphabet and Swahili word analogy dataset to address the need for language processing resources especially for low resource languages. Therefore, we derive the unannotated Swahili dataset by pre-processing raw Swahili data using a Python script, formulate the syllabic alphabet and develop the Swahili word analogy dataset based on an existing English dataset. We envisage that the datasets will not only support language models but also other NLP downstream tasks such as part-of-speech tagging, machine translation and sentiment analysis.

Visit

www.sciencedirect.compubmed.ncbi.nlm.nih.gov

Connected records

dataset

Tasks

language modelingembeddings

Languages

Swahili

Licenses

Attribution 4.0 International

Similaires

Low-Resource Language Modelling of South African LanguagesLungisanikhan/Enhancing-Semantic-Relatedness-for-Low-Resource-African-Languages-Insights into Low-Resource Language Modelling: Improving Model Performances for South African LanguagesOvercoming Data Scarcity in Generative Language Modelling for Low-Resource Languages: A Systematic ReviewEnhancing Sentiment Analysis in Amharic: Leveraging Transformer-Based Language Model for Low-Resource African LanguagesContinual-learning for Modelling Low-Resource Languages from Large Language Models

Low-Resource Language Modelling of South African Languages

Language models are the foundation of current neural network-based models for natural language under

Lungisanikhan/Enhancing-Semantic-Relatedness-for-Low-Resource-African-Languages-

Enhancing Semantic Relatedness for Low-Resource African Languages via Transfer Learning and M2M-100

Insights into Low-Resource Language Modelling: Improving Model Performances for South African Languages

To address the gap in natural language processing for Southern African languages, our paper presents

Overcoming Data Scarcity in Generative Language Modelling for Low-Resource Languages: A Systematic Review

Generative language modelling has surged in popularity with the emergence of services such as ChatGP

Enhancing Sentiment Analysis in Amharic: Leveraging Transformer-Based Language Model for Low-Resource African Languages

Continual-learning for Modelling Low-Resource Languages from Large Language Models

Modelling a language model for a multi-lingual scenario includes several potential challenges, among