Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Enhancing Multi-Label News Text Classification for an Understudied Language: A Comprehensive Study on CNN Performance and Pre-Trained Word Embeddings

Domain:

natural language processing

Record type:

dataset
Creator:
DirAru
Publisher:
Uni
Host:
Today's news texts are classified using a multi-label system, which allows for the assignment of a potentially large number of labels to specific instances. The majority of earlier scholars have only looked into mutual exclusion at a single level. Nonetheless, the primary goal of this study was to categorise the news material using multiple labels. Many text documents are created these days from a variety of offline and internet sources. This generated news text is disordered state. As a result, timely access to the needed content from the sources is challenging. Compared with traditional text classification, multi-label classification is difficult and challenging because of its multi-dimensional labels. Convolutional neural networks are used in this study's tests on the problem domain for Afaan Oromo multi-label news text classification due to their ease of assimilation of pre-trained word embeddings. According to pre-trained word embedding with a train-test ratio of 10/90, the new proposed model has shown improved performance. The suggested CNN models might be helpful for labelling news articles in Afaan Oromo news text. The goal of many researchers working on Afaan Oromo classifier development is to use various learning algorithms to boost classification accuracy as the number of categories or labels increases. Using various approaches, they attempted to use basic machine learning methods to address the calculation time issue. Unfortunately, all low-resource language researchers focus on flat, hierarchical, and multi-class classification types, but we created a model for multi-label text classification and attempted to apply it using a deep learning algorithm. Over 5640 Afaan Oromo news dataset items are analysed experimentally over eight main news categories. Python served as our experimental platform for both text classification and word embedding. After the model is fully implemented, the best result of the precision, recall, F1 score and accuracy rate train test ratio of 10/90 for pertained word_ embedding is 89.7%, 88.6%,  93.3% and 96.5, respectively.

Visit

doi.org

Tasks

news classificationtext classificationtopic classification

Languages

OromoOromo, Borana-Arsi-Guji

Licenses

https://creativecommons.org/licenses/by/4.0

Similar

Evaluating Performance of Pre-trained Word Embeddings on Assamese, a Low-resource LanguageEnhancing Natural Language Processing in Somali Text Classification: A Comprehensive Framework for Stop Word RemovalIntent Classification Using Pre-trained Language Agnostic Embeddings For Low Resource LanguagesAFAAN OROMO MULTI LABEL NEWS TEXT CLASSIFICATION USING DEEP LEARNING APPROACHAUTOMATIC AFAAN OROMO MULTI LABEL NEWS TEXT CLASSIFICATION USING NEURAL NETWORK APPROACHAnalyzing Acoustic Word Embeddings from Pre-trained Self-supervised Models

Evaluating Performance of Pre-trained Word Embeddings on Assamese, a Low-resource Language

Enhancing Natural Language Processing in Somali Text Classification: A Comprehensive Framework for Stop Word Removal

Intent Classification Using Pre-trained Language Agnostic Embeddings For Low Resource Languages

Building Spoken Language Understanding (SLU) systems that do not rely on language specific Automatic

AFAAN OROMO MULTI LABEL NEWS TEXT CLASSIFICATION USING DEEP LEARNING APPROACH

The development of the internet has made Afaan Oromo's writings widely available both offline and on

AUTOMATIC AFAAN OROMO MULTI LABEL NEWS TEXT CLASSIFICATION USING NEURAL NETWORK APPROACH

Major Advisor: - GetachewMamo(PHD) The classification of natural language texts has gained a growin

Analyzing Acoustic Word Embeddings from Pre-trained Self-supervised Models

IEEE ICASSP 2023 Conference, Hybrid Event, 4-10 June 2023, Rhodes Island, Greece Given the strong re