Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Better Word Representation Vectors Using Syllabic Alphabet: A Case Study of Swahili

Domain:

natural language processing

Record type:

datasetmodel
Creator:
CasZhoLiuRef
Publisher:
MDP
Host:
Deep learning has extensively been used in natural language processing with sub-word representation vectors playing a critical role. However, this cannot be said of Swahili, which is a low resource and widely spoken language in East and Central Africa. This study proposed novel word embeddings from syllable embeddings (WEFSE) for Swahili to address the concern of word representation for agglutinative and syllabic-based languages. Inspired by the learning methodology of Swahili in beginner classes, we encoded respective syllables instead of characters, character n-grams or morphemes of words and generated quality word embeddings using a convolutional neural network. The quality of WEFSE was demonstrated by the state-of-art results in the syllable-aware language model on both the small dataset (31.229 perplexity value) and the medium dataset (45.859 perplexity value), outperforming character-aware language models. We further evaluated the word embeddings using word analogy task. To the best of our knowledge, syllabic alphabets have not been used to compose the word representation vectors. Therefore, the main contributions of the study are a syllabic alphabet, WEFSE, a syllabic-aware language model and a word analogy dataset for Swahili.

Visit

doi.org

Tasks

embeddingslanguage modeling

Languages

Swahili

Licenses

https://creativecommons.org/licenses/by/4.0/

Similar

Learning Syllables Using Conv-LSTM Model for Swahili Word Representation and Part-of-speech TaggingfastText Pretrained Word Vectorsjordiaphane/amharic-word-vectorsA Word Game Support Tool Case StudyLearning Word Vectors for 157 LanguagesZeinab-Haroon/Exploring-Word-Vectors-NLP

Learning Syllables Using Conv-LSTM Model for Swahili Word Representation and Part-of-speech Tagging

The need to capture intra-word information in natural language processing (NLP) tasks has inspired r

fastText Pretrained Word Vectors

We distribute pre-trained word vectors for 157 languages, trained on Common Crawl and Wikipedia using fastText. These models were trained using CBOW with position-weights, in dimension 300, with character n-grams of length 5, a window of size 5 and 10 negatives. We

jordiaphane/amharic-word-vectors

fasttext amharic bin loaded to word and vector dicts # amharic-word-mappings fasttext amharic bin l

A Word Game Support Tool Case Study

International audience This article reports on the approach taken, experience gathere

Learning Word Vectors for 157 Languages

Distributed word representations, or word vectors, have recently been applied to many tasks in natural language processing, leading to state-of-the-art performance. A key ingredient to the successful application of these representations is to train them on very lar

Zeinab-Haroon/Exploring-Word-Vectors-NLP

An assignment that was given as part of the 'Exploring Word vectors-NLP Course' at the African Maste