Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

TRANSFORMER-BASED SPAM DETECTION FOR UZBEK LANGUAGE: A COMPARATIVE STUDY WITH TRADITIONAL MACHINE LEARNING MODELS

Domain:

natural language processing

Record type:

paper
Creator:
Ata
Publisher:
Zenodo
Host:avatar
Spam detection remains a critical challenge in natural language processing, particularly for low-resource languages such as Uzbek. While traditional machine learning approaches have been widely applied to text classification tasks, their reliance on handcrafted features limits contextual understanding. Recent advances in Transformer-based architectures, especially BERT (Bidirectional Encoder Representations from Transformers), have demonstrated superior performance in capturing semantic relationships within text. This study proposes a BERT-based spam detection model for Uzbek SMS messages and compares its effectiveness with conventional machine learning models including Naïve Bayes, Support Vector Machine (SVM), and Logistic Regression. A labeled Uzbek SMS dataset was utilized, divided into training and testing subsets. Traditional models were trained using TF-IDF feature extraction, while a multilingual BERT model was fine-tuned for binary classification. Experimental results indicate that the Transformer-based model significantly outperforms classical approaches in terms of accuracy, precision, recall, and F1-score. The findings confirm that contextual embeddings are highly effective for spam detection in morphologically rich and low-resource languages. The research contributes to Uzbek NLP by providing one of the first systematic comparative analyses between deep contextual models and traditional machine learning techniques for spam classification. Practical implications include the potential integration of the proposed model into mobile communication platforms. Limitations include dataset size and computational requirements, which future studies may address using lightweight Transformer architectures.

Visit

doi.orgzenodo.org

Tasks

text classification

Licenses

Creative Commons Attribution 4.0 Internationalhttps://creativecommons.org/licenses/by/4.0/legalcode

Similar

Comparative Analysis of Machine Learning Models for Detecting Mobile Messaging Spam In Swahili SMSPolyTruth: Multilingual Disinformation Detection using Transformer-Based Language ModelsAdvancing Amharic Information Retrieval: A Comparative Analysis of Traditional, Neural, and Transformer-Based ModelsA Comparative Study of Transformer-based Models for Hate-Speech Detection in English-Kiswahili Code-Switched Social Media TextA Comparative Study of Embedding Models for Offensive Language Detection in NigeriaComparative Analysis of Deep Learning and Transformer Models for Low Resource Hausa Language Text Summarization

Comparative Analysis of Machine Learning Models for Detecting Mobile Messaging Spam In Swahili SMS

PolyTruth: Multilingual Disinformation Detection using Transformer-Based Language Models

Disinformation spreads rapidly across linguistic boundaries, yet most AI models are still benchmarke

Advancing Amharic Information Retrieval: A Comparative Analysis of Traditional, Neural, and Transformer-Based Models

A Comparative Study of Transformer-based Models for Hate-Speech Detection in English-Kiswahili Code-Switched Social Media Text

The transformer architecture, first introduced in 2017 by researchers at Google, has revolutionized

A Comparative Study of Embedding Models for Offensive Language Detection in Nigeria

Offensive language detection has become a pivotal challenge in Natural Language Processing (NLP), pa

Comparative Analysis of Deep Learning and Transformer Models for Low Resource Hausa Language Text Summarization

Natural language processing for low-resource languages presents unique and multifaceted challenges t