Abstract:
Despite the emergence of large-scale multilingual pre-trained models like mBERT, XLM-RoBERTa, and mT5, natural language processing (NLP) still struggles in low-resource languages due to limited annotated data. This paper explores the use of transfer learning to adapt pre-trained multilingual models to low-resource tasks such as Named Entity Recognition (NER), sentiment analysis, and machine translation for languages like Amharic, Hausa, and Sinhala. By leveraging zero-shot and few-shot learning paradigms and evaluating cross-lingual embeddings and token overlap, we demonstrate significant improvements in model performance. Regression analysis confirms the predictive value of embedding similarity and token overlap, and SHAP-based interpretability reveals transparent model behaviors.
Keywords:
Multilingual NLP, Low-Resource Languages, Transfer Learning, mBERT, XLM-RoBERTa, mT5, Few-Shot Learning, Cross-Lingual Embeddings, SHAP, LIME