Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Author identification for Under-Resourced language (KadazanDusun)

Domaine:

natural language processing

Type de record:

paper
Créateur:
NurSuhDay
Éditeur:
Institute of Advanced Engineering and Science
Hôte:
This paper presents the task of Author Identification for KadazanDusun language by using tweets as the source of data to perform Author Identification task of short text on KadazanDusun, which is considered as one the under-resourced language in Malaysia. The aim of this paper is to demonstrate Author Identification of short text on KadazanDusun. Besides, this paper also examines the performance of two machine learning algorithms on the KadazanDusun data set by analyzing the stylometric features. Stylometric features are used to quantify the writing styles of the authors which includes character n-grams and word n-grams. The workflow of Author Identification implements the machine learning approach to solve the single-labelled multi-class problem and predict the author of a given message in KadazanDusun. Two classifiers are used to compare the accuracy including Naïve Bayes and Support Vector Machine (SVM). The results show that the combination of n-grams which is word-level unigram and {1-5}-grams with character 3-grams are the most relevant stylometric features in identifying the author of KadazanDusun message with an accuracy of 80.17%. The results also show that SVM classifier has outperformed Naive Bayes in this Author Identification task with the accuracy of 80.17%.

Visit

doi.org

Tasks

text classification

Licenses

http://creativecommons.org/licenses/by-nc/4.0

Similaires

The Application of Computer-Aided Under-Resourced Language Translation for Malay into KadazandusunShort Text Language Identification for Under Resourced LanguagesTOWARDS CURBING CYBER-BULLYING IN MALAYSIA BY AUTHOR IDENTIFICATION OF IBAN AND KADAZANDUSUN OSN TEXT USING DEEP LEARNINGibrahimbukhari1998/-Zero-Shot-for-Under-Resourced-LanguageBenchmarking Multi-Task Learning for Sentiment Analysis and Offensive Language Identification in Under-Resourced Dravidian LanguagesDocument Classification for the Under-resourced Amharic Language

The Application of Computer-Aided Under-Resourced Language Translation for Malay into Kadazandusun

A computer-aided language translation using a Machine translation (MT) is an application performed b

Short Text Language Identification for Under Resourced Languages

The paper presents a hierarchical naive Bayesian and lexicon based classifier for short text languag

TOWARDS CURBING CYBER-BULLYING IN MALAYSIA BY AUTHOR IDENTIFICATION OF IBAN AND KADAZANDUSUN OSN TEXT USING DEEP LEARNING

Online Social Network (OSN) is frequently used to carry out cyber-criminal actions such as cyberbull

ibrahimbukhari1998/-Zero-Shot-for-Under-Resourced-Language

Zero-Shot POS Tagging for Under-Resourced Languages # Cross-Lingual POS Tagging: XLM-R vs. Glot500

Benchmarking Multi-Task Learning for Sentiment Analysis and Offensive Language Identification in Under-Resourced Dravidian Languages

To obtain extensive annotated data for under-resourced languages is challenging, so in this research

Document Classification for the Under-resourced Amharic Language

NLP is severely hampered by a scarcity of digital resources. This is especially true for Amharic, a