Logo Lanfrica

CO-REFERENCE RESOLUTION MODEL FOR AFAAN OROMO TEXT: KNOWLEDGE POOR APPROACH

Domaine:

natural language processing

Type de record:

modelpaper
Créateur:
ABD
Éditeur:
Zenodo
Hôte:avatar
Kamal Mohammed (Assistant Professor) Major Advisor Linking different expressions in a text to the same real-world entity is known as co-reference resolution, and it is an essential component of natural language understanding. Designing a knowledge-poor co-reference resolution model for Afaan Oromo text is the primary goal of this study. The distinctive linguistic characteristics of Afaan Oromo have not been adequately covered by the research that is currently available, making it difficult to recognize and resolve co-reference in its textual contexts. This issue limits Afaan Oromo's usefulness in domains like information extraction, machine translation, and text summarization and impedes the development of natural language processing applications for the language. By creating a knowledge-poor co-reference resolution model especially suited to the intricacies of Afaan Oromo language structure and usage, this study seeks to close this gap. Data collection, annotation, preprocessing, and the use of deep learning techniques for model training and evaluation are all included in the methodology. We used Afaan Oromo text from several reliable sources for this investigation. e. AO textbooks, AO news, and AO Bible. Label encoding, normalization, and cleaning are all part of the pre-processing phase. Additionally, the train_test splitter method is used to separate the preprocessed dataset into training data (eighty percent of the original data) and testing data (20 percent of the original data). Model construction, compilation, and training are all included in the training phase. In the testing phase, the model is also evaluated and saved. We use the AO dataset, which comprises 5614 rows of data prepared from multiple trustworthy sources for this study, to test the suggested model. Accuracy, precision, recall, and F-measure are the evaluation metrics most frequently used for co-reference resolution tasks. With performance metrics of Precision: 86%, Recall: 74% percent, F1 Score: 71.9% and Accuracy: 87%, the co-reference resolution model created for Afaan Oromo texts produced encouraging results. The findings show that the co-reference resolution model is not only efficient but also able to support a range of NLP applications, including machine translation and information extraction. Future research should concentrate on enhancing recall even more, possibly by incorporating more sophisticated linguistic features and a wider variety of training data. Key Words: Afaan Oromo, Co-reference resolution, Knowledge poor approach, Natural Language Preprocessing