Logo Lanfrica

QUERY ORIENTED AFAAN OROMO TEXT SUMMARIZATION USING SENTENCE FEATURE EXTRACTION

Domaine:

natural language processing
Créateur:
Neg
Éditeur:
Zenodo
Hôte:avatar
Major Advisor: Kemal Mohammed (Assist. Professor) ABSTRACT In the recent years, information grows rapidly along with the development of social media. With the increasing amount of information, it takes more effort and time to review the entire text document and understand its contents. One possible solution to the above problem is to read the summary of the document. The summary will not only retain the essence of the document, but will also save a lot of time and effort. An effective summary of the document will concise and fluent while preserving key information and overall meaning. Automatic text summarizer is one of the various tools used for the purpose of shortening lengthy documents, and alleviating the type of problem. This thesis focuses on developing query oriented Afaan Oromo text summarizer, through systematic integration of features: sentence position, keyword frequency, cue phrase, sentence length handler, part of speech tagging(POS), pronoun, named entities, occurrence of numbers and events in sentences and unigram, bigram words are focused in this thesis. In addition to this prototype user interface are developed for the purpose of testing text summary. For this task, we have used Vector Space Model (VSM) and the cosine similarity measure to find the most relevant sentence extracted for Afaan Oromo text summary. In this thesis Recall, Precision and F-measure are used as evaluation metrics using ROUGE tools. Being trained and tested on the dataset of size 6032 tokens, for validation and testing 20 different document topics are collected. From this 15 of them have been used for training. while, the rest 5 documents are used for testing. QOAOTS performed using ROUGE with extraction rate of 20%, 25% and 30%. The system performed the results for 25% extraction rates, 84% Precision, 97.5% Recall and 89.5% F measure.