Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Multilingual Cyber Threat Intelligence Vendor Classification Using Multilingual BERT (mBERT): An Empirical Analysis of Class Imbalance Mitigation Strategies

Domaine:

natural language processing

Type de record:

paper
Créateur:
MarIsmMohKab
Éditeur:
Elsevier BV
Hôte:
When a critical zero-day vulnerability is disclosed in a French-language security advisory or a Hausa-language threat bulletin, most deployed vendor identification systems simply fail, not because the vulnerability is obscure, but because the intelligence pipeline was never designed to read anything but English. We present an empirical evaluation of a multilingual vendor identification system for Cyber Threat Intelligence (CTI) that applies contextualised embeddings from Multilingual BERT (mBERT, bert-base-multilingual-cased) to classify vulnerability text across five languages; English, French, Hausa, Yoruba, and Igbo into 174 distinct vendor categories drawn from the NIST National Vulnerability Database (NVD). The AI contribution is a systematic controlled ablation of three class imbalance mitigation strategies; Synthetic Minority Over-sampling Technique (SMOTE), weighted cross-entropy loss, and their combination applied to a custom ThreatClassifier neural network head trained on frozen mBERT embeddings. The engineering application is automated vendor attribution in operational CTI pipelines, where accurate and fast vendor identification directly governs patch prioritisation . Results show that SMOTE oversampling improves macro F1 by 5.3% over the unmitigated baseline (0.8938 vs. 0.8409), while weighted loss alone reduces overall accuracy by 32% under extreme 174-class imbalance, a finding that contradicts common practice. The hybrid strategy combining SMOTE with weighted loss produces results statistically identical to SMOTE alone, exposing a fundamental redundancy: SMOTE’s rebalancing of the training distribution neutralises the corrective mechanism that weighted loss was designed to provide. Precision-Recall Average Precision reaches 0.977 (micro) and 0.979 (macro) for the best model, against a random baseline

Visit

doi.org

Tasks

text classification

Languages

HausaYoruba

Licenses

https://www.uspto.gov/ip-policy/copyright-policy/copyright-basics

Similaires

Predicting Personality from Social Media Text Using Multilingual BERT (mBERT)Sentiment Classification in Swahili Language Using Multilingual BERTCyber Threat Mitigation Strategies for Financial Systems in East Africa: A Quantitative ApproachCyber Threat Intelligence : Challenges and OpportunitiesSpecializing Multilingual Language Models: An Empirical StudyAutomatic classification of multilingual occupations using transfer learning

Predicting Personality from Social Media Text Using Multilingual BERT (mBERT)

The rapid growth of social media platforms has generated an enormous volume of user-generated textua

Sentiment Classification in Swahili Language Using Multilingual BERT

The evolution of the Internet has increased the amount of information that is expressed by people on

Cyber Threat Mitigation Strategies for Financial Systems in East Africa: A Quantitative Approach

Cyber threats to financial systems in East Africa pose significant risks to economic stabil

Cyber Threat Intelligence : Challenges and Opportunities

The ever increasing number of cyber attacks requires the cyber security and forensic specialists to

Specializing Multilingual Language Models: An Empirical Study

Pretrained multilingual language models have become a common tool in transferring NLP capabilities t

Automatic classification of multilingual occupations using transfer learning

Socio-economic research often requires detailed information on individual occupations in order to st