Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Data-Driven Part-of-Speech Tagging for the Gikuyu Language: Development, Challenges, and Prospects

Domain:

natural language processing

Record type:

paper
Creator:
Gab
Publisher:
Aca
Host:
This paper presents the development of a data-driven Part-of-Speech (POS) tagger for Gikuyu, a Bantu language spoken in Kenya. Gikuyu, like many indigenous African languages, is under-resourced, with limited computational tools for linguistic processing. By employing a corpus sourced primarily from the Gikuyu Bible and leveraging a Memory-Based Tagging (MBT) approach, this study demonstrates the feasibility of creating a robust POS tagging system. The tagger achieved a precision of 90.44%, a recall of 88.34%, and an F-score of 91.35%. These results underscore its potential for applications in machine translation, speech recognition, and language preservation. The study highlights the challenges of working with under-resourced languages, including data collection and annotation, and provides recommendations for future work, including integration with broader NLP tasks.

Visit

doi.org

Tasks

part of speech tagging

Languages

Gikuyu

Similar

DIO:10.5121/ijnlc.2024.13602 15 DATA-DRIVEN PART-OF-SPEECH TAGGING FOR THE GIKUYU LANGUAGE: DEVELOPMENT, CHALLENGES AND PROSPECTSData-Driven Part-of-Speech Tagging of KiswahiliPart of Speech Tagging for Setswana African Language

DIO:10.5121/ijnlc.2024.13602 15 DATA-DRIVEN PART-OF-SPEECH TAGGING FOR THE GIKUYU LANGUAGE: DEVELOPMENT, CHALLENGES AND PROSPECTS

Data-Driven Part-of-Speech Tagging for the Gikuyu Language: Development, Challenges, and Prospects

Data-Driven Part-of-Speech Tagging of Kiswahili

Part of Speech Tagging for Setswana African Language