# Natural Language Processing (NLP) Research on Swahili Language
**Swahili**, also **Kiswahili**, is a Bantu language and the native language of the Swahili people. It is spoken by 120 – 150 million people across East and Southern Africa. Originates from Bantu languages of the coast of East Africa (Pokomo, Taita and Mijikenda languages etc.). Opinions vary, however, about 20% of the Swahili vocabulary is derived from loan words. The vast majority Arabic, but also other contributing languages, including Persian, Hindustani, Portuguese, German & Malay. It is an official language in Kenya, Tanzania, Uganda & Rwanda and is recognized as a minority language in DRC, Burundi & Mozambique.
This work is not a masterpiece however, it entails related works revolving around NLP research on Swahili language including; NLP tasks on Swahili Language, available datasets on Swahili Language and NLP Models trained on Swahili language.
# i) NLP Tasks on Swahili Language
## Parts-of-speech tagging (POST)
**Papers**
[[Paper]](dl.acm.org) Shikali, et al. (2021) _Learning Syllables Using Conv-LSTM Model for Swahili Word Representation and Part-of-speech Tagging, ACM._
[[Paper]](biblio.ugent.be) Guy, et al. (2012) _Resource-Light Bantu Part-of-Speech Tagging, Google Scholar._
**Available datasets**
[[Dataset]](kielipankki.fi) _Helsinki Corpus of Swahili 2.0._
## Word Sense Disambiguation (WSD)
**Papers**
[[Paper]](helda.helsinki.fi) Wanjiku NG’ANG’A (2005) _Word Sense Disambiguation of Swahili: Extending Swahili Language Technology with Machine Learning, Google Scholar._
## Named Entity Recognition (NER)
**Papers**
[[Paper]](MasakhaNER: Named Entity Re…) Adelani, et al. (2021) _MasakhaNER: Named Entity Recognition for African Languages._
**Available datasets**
[[Dataset]](github.com …