Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

AUTHORSHIP ATTRIBUTION MODEL FOR AFAAN OROMO DOCUMENTS OF SOCIAL MEDIA USING DEEP LEARNING APPROACH

Domaine:

natural language processing

Type de record:

paper
Créateur:
Gir
Éditeur:
Zenodo
Hôte:avatar
Major Advisor: Mr. Kamal Mohammed (Ass. Professor) In today's digital world, where we share and talk using technology, uncovering the secrets hidden in writing is really important. To do so Looking into authorship attribution can be very interesting. This study focuses on creating Authorship Attribution (AA) model for Afaan Oromo documents. The aim is to classify and predict the authorship of an unidentified text from a predetermined group of potential authors or attributing writings of unknown origin to one of possible writers. The aim of this study is to utilize deep learning approaches with word embedding on authorship attribution for Afaan Oromoo documents of social media using LSTM, BiLSTM and CNN algorithms and to recommend the best model for AA problems for Afaan Oromoo documents. To develop this model we collected a total of 10,485 Facebook posts from verified Facebook pages that employ influence in politics, the economy, religion, culture and individuals who possess the ability to easily engage the community and guide them toward shared objectives. The dataset was initially divided into training and testing sets at an 80/20% and 70/30 % ratio. This allocation meant that 80% of the data (8388 data) was designated for training, while the remaining 20% (2097 data) was reserved for testing. Additionally, in another split with a 70/30% ratio, 7339 data were assigned to training, and 3146 data were assigned to testing. Furthermore, we set aside 10% of the training sample for validation. In this work, different Natural language processing tasks such as text preprocessing which include tokenization, stemming, lemmatization, stop words removal and normalization are performed. Finally, the result of our models is compared. We compared LSTM, BILSTM and CNN with the same dataset and same parameters. However, we obtained different performance on 80/20% and 70/30% ratio. The model has high accuracy on 80/20%. BiLSTM has a great performance with 92% accuracy and 88% precision and LSTM and CNN have got 91% accuracy and 90% precision and 91% accuracy and 89% precision respectively. 

Visit

doi.orgzenodo.org

Tasks

text classification

Languages

OromoOromo, Borana-Arsi-GujiOromo, EasternOromo, West Central

Licenses

Open Data Commons Attribution Licensehttp://www.opendefinition.org/licenses/odc-byOpen Accessinfo:eu-repo/semantics/openAccess