Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

TRANSFORMER-BASED SPAM DETECTION FOR UZBEK LANGUAGE: A COMPARATIVE STUDY WITH TRADITIONAL MACHINE LEARNING MODELS

Domaine:

natural language processing

Type de record:

paper
Créateur:
Ata
Éditeur:
Zenodo
Hôte:avatar
Spam detection remains a critical challenge in natural language processing, particularly for low-resource languages such as Uzbek. While traditional machine learning approaches have been widely applied to text classification tasks, their reliance on handcrafted features limits contextual understanding. Recent advances in Transformer-based architectures, especially BERT (Bidirectional Encoder Representations from Transformers), have demonstrated superior performance in capturing semantic relationships within text. This study proposes a BERT-based spam detection model for Uzbek SMS messages and compares its effectiveness with conventional machine learning models including Naïve Bayes, Support Vector Machine (SVM), and Logistic Regression. A labeled Uzbek SMS dataset was utilized, divided into training and testing subsets. Traditional models were trained using TF-IDF feature extraction, while a multilingual BERT model was fine-tuned for binary classification. Experimental results indicate that the Transformer-based model significantly outperforms classical approaches in terms of accuracy, precision, recall, and F1-score. The findings confirm that contextual embeddings are highly effective for spam detection in morphologically rich and low-resource languages. The research contributes to Uzbek NLP by providing one of the first systematic comparative analyses between deep contextual models and traditional machine learning techniques for spam classification. Practical implications include the potential integration of the proposed model into mobile communication platforms. Limitations include dataset size and computational requirements, which future studies may address using lightweight Transformer architectures.

Visit

doi.orgzenodo.org

Tasks

text classification

Licenses

Creative Commons Attribution 4.0 Internationalhttps://creativecommons.org/licenses/by/4.0/legalcode