Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

A topic modeling approach for analyzing and categorizing electronic healthcare documents in Afaan Oromo without label information

Domain:

natural language processinghealthcare

Record type:

paper
Creator:
EtaMriTek
Publisher:
Spr
Host:
Abstract Afaan Oromo is a resource-scarce language with limited tools developed for its processing, posing significant challenges for natural language tasks. The tools designed for English do not work efficiently for Afaan Oromo due to the linguistic differences and lack of well-structured resources. To address this challenge, this work proposes a topic modeling framework for unstructured health-related documents in Afaan Oromo using latent dirichlet allocation (LDA) algorithms. All collected documents lack label information, which poses significant challenges for categorizing the documents and applying the supervised learning methods. So, we utilize the LDA model since it offers solutions to this problem by allowing discovery of the latent topics of the documents without requiring the predefined labels. The model takes a word dictionary to extract hidden topics by evaluating word patterns and distributions across the dataset. Then it extracts the most relevant document topics and generates weight values for each word in the documents per topic. Next, we classify the topics using the represented keyword as input and assign class labels based on human evaluations topic coherence. This model could be applied to classifying medical documents and used to find specialists who best suitable for patients’ requests from the obtained information. As a conclusion of our findings, the topic modeling using LDA gave the promised value of 79.17% accuracy and 79.66% F1 score for test documents of the dataset.

Visit

doi.org

Tasks

text classificationtopic classification

Languages

OromoOromo, Borana-Arsi-Guji

Licenses

https://creativecommons.org/licenses/by/4.0https://creativecommons.org/licenses/by/4.0

Similar

AFAAN OROMO SHORT TEXT CLUSTERING USING TOPIC MODELING IN DEEP LEARNING APPROACHIntelligence Information Retrieval System Modeling for Afaan OromoAUTOMATIC MULTI LABEL CLASSIFICATION FOR AFAAN OROMO DOCUMENT USING DEEP LEARNING APPROACHAFAAN OROMO MULTI LABEL NEWS TEXT CLASSIFICATION USING DEEP LEARNING APPROACHAUTHORSHIP ATTRIBUTION MODEL FOR AFAAN OROMO DOCUMENTS OF SOCIAL MEDIA USING DEEP LEARNING APPROACHAUTOMATIC AFAAN OROMO MULTI LABEL NEWS TEXT CLASSIFICATION USING NEURAL NETWORK APPROACH

AFAAN OROMO SHORT TEXT CLUSTERING USING TOPIC MODELING IN DEEP LEARNING APPROACH

Main Advisor: Mr. Wakgari Dibaba (Ass.Prof ) Text clustering groups texts according to a particular

Intelligence Information Retrieval System Modeling for Afaan Oromo

AUTOMATIC MULTI LABEL CLASSIFICATION FOR AFAAN OROMO DOCUMENT USING DEEP LEARNING APPROACH

Advisor: Mr. Kasahun Abdisa (Phd candidate) Document classification is a technique which classifies

AFAAN OROMO MULTI LABEL NEWS TEXT CLASSIFICATION USING DEEP LEARNING APPROACH

The development of the internet has made Afaan Oromo's writings widely available both offline and on

AUTHORSHIP ATTRIBUTION MODEL FOR AFAAN OROMO DOCUMENTS OF SOCIAL MEDIA USING DEEP LEARNING APPROACH

Major Advisor: Mr. Kamal Mohammed (Ass. Professor) In today's digital world, where we share and tal

AUTOMATIC AFAAN OROMO MULTI LABEL NEWS TEXT CLASSIFICATION USING NEURAL NETWORK APPROACH

Major Advisor: - GetachewMamo(PHD) The classification of natural language texts has gained a growin