Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

HIDDEN MARKOV MODEL BASED AFAAN OROMO NAMED ENTITY RECOGNITION

Domain:

natural language processing

Record type:

paper
Creator:
ADU
Publisher:
Zenodo
Host:avatar
Major Advisor: GETACHEW MAMO (PhD) Named Entity Recognition (NER) is a subtask of Natural Language Processing (NLP) which plays a vital role to achieve human level performance on specific documents such as news texts to identify and classify named entities. The purpose of this research is to model NER system which extracts NEs from Afaan Oromo news texts which can help to identify events of specified things in running text using Hidden Markov Model (HMM) approach. For this study, out of 32,522 tokens we gathered for our work, we used 75% of it, which counted to more than 26,000 tokens to train the model. Thus, after the model has been created, we could see that the training corpus consisted tokens of which 578 of them represent person names, 1,146 of them represent organization names, 815 of them represent location names, 357 of them represent measurements, 309 of them represent date and time expressions, 493 of them represent miscellaneous names, and 22,324 of them were not named entities. Likewise, we used 25% of it, which counted to about 6,500 separate tokens to test whether the model properly identify and recognize NE tags of these tokens. For implementation of the prototype of the system, we used the most popular NLP supporting language known as Python programming language along with some of its built-in modules. At the end, we could experiment how actually the size of our training data would affect the performance of the system as a whole. To do so, we divided our total training dataset into two: 60% of it, which counted to 15,613 tokens, as training data and the other 40% of it, which counted to 10,409 tokens, as experimentation data. Finally, we could observe that accuracy of the system was become well (i.e. which showed an increase in about 6.81% F-score) when we used the combination of these datasets together as training data than using only the original training data. And also, we used ten-fold cross validation method in which the whole dataset is randomly partitioned into two; i.e., 90% of it is used for training purpose and the other 10% is used for testing the performance of the system. In this case, such partitioning process takes place for ten times and at each step the accuracy result is taken so that at the end of this process the average result of all steps is taken. Thus, according to this method, the accuracy result obtained was Precision of 92.78%, Recall of 92.83% and F-Score of 92.80%. Keywords: Afaan Oromo, Natural Language Processing, Named Entities, Named Entity Recognition, Hidden Markov Model

Visit

doi.orgzenodo.org

Tasks

information extractionnamed entity recognition

Languages

OromoOromo, Borana-Arsi-Guji

Licenses

Creative Commons Attributionhttp://www.opendefinition.org/licenses/cc-byOpen Accessinfo:eu-repo/semantics/openAccess