Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

A Basic Language Resource Kit Implementation for the Igbo <i>NLP</i> Project

Domain:

natural language processing

Record type:

datasetpaper
Creator:
IkeMarUchIgn
Publisher:
Ass
Host:
Igbo, an African language with around 32 million speakers worldwide, is one of the many languages having few or none of the language processing resources needed for advanced language technology applications. In this article, we describe the approach taken to creating an initial set of resources for Igbo, including an electronic text corpus, a part-of-speech (POS) tagset, and a POS-tagged subcorpus. We discuss the approach taken in gathering texts, the preprocessing of these texts, and the development of the POS tagged corpus. We also discuss some of the problems encountered during corpus and tagset development and the solutions arrived at for these problems.

Visit

doi.org

Tasks

part of speech tagging

Languages

Igbo

Licenses

https://www.acm.org/publications/policies/copyright_policy#Background