Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

AFAAN OROMO SHORT TEXT CLUSTERING USING TOPIC MODELING IN DEEP LEARNING APPROACH

Domain:

natural language processing
Creator:
BY
Publisher:
Zenodo
Host:avatar
Main Advisor: Mr. Wakgari Dibaba (Ass.Prof ) Text clustering groups texts according to a particular characteristic to determine how similar they are to one another. Keyword-based models have been used as a feature recently, such as the TFIDF model for texts. A key-word-based approach is not feasible for short texts due to their briefness. Moreover, it lacks semantic structure, which limits the range of further text analysis. Moreover, it lacks semantic structure, which limits the range of further text analysis. The topic model was developed to determine the probability distributions of subjects across a predefined collection of phrases or language. Unlike the TFIDF, the topic model has a semantic structure of texts. In addition to ids, the topic model can cluster according to the topic of the cluster. In order to use deep learning to uncover hidden subjects in a collection of short texts, we used topic modeling in this thesis. These days, BERTopic Modeling is one of the most popular and often used topic modeling strategies for deep learning. The main goal of this thesis study is to develop a model for short Afaan Oromo text clustering using the topic modeling approach. We used the BERTopic algorithm as a topic modeling tool. BERTopic enhances topic modeling by using topic based embedding to capture the semantic meaning of text and clustering methods to identify logical subjects. Because it offers powerful visualization tools together with dynamic and hierarchical topic modeling capabilities, it is a versatile and effective method for locating and analyzing topics in text data. To improve topic extraction and clustering accuracy, neural word embedding and topic models have been merged. We have concluded that six clusters is the optimal amount for our experiment based on the Average Coherence Method. We employed 4639 Afan Oromo documents in our training experiment, which covered subjects including politics, health, sports, science, technology, and peace. In terms of word embedding as feature extraction, the final result shows that the model's overall accuracy is 79.24%. Last but not least, we evaluated the BERTopic model using the Coherence score, and we received an 81.34% score. This implies: Semantic similarities between words within themes are more substantial and comprehensible.
Key Words: BERTopic Modeling, Text Clustering, Coherence Model, Recurrent Neural Network (RNN)

Visit

doi.orgzenodo.org

Tasks

topic classificationtext classification

Languages

OromoOromo, Borana-Arsi-GujiOromo, EasternOromo, West Central

Licenses

Open Group Test Suite Licensehttp://www.opengroup.org/testing/downloads/The_Open_Group_TSL.txtOpen Accessinfo:eu-repo/semantics/openAccess