Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Lungisanikhan/Enhancing-Semantic-Relatedness-for-Low-Resource-African-Languages-

Domaine:

natural language processing

Type de record:

project
Créateur:
Lun
Hôte:
Enhancing Semantic Relatedness for Low-Resource African Languages via Transfer Learning and M2M-100 Data Augmentation # Semantic Relatedness for Low-Resource African Languages This repository contains the implementation and results for the project: **_Enhancing Semantic Relatedness for Low-Resource African Languages via Transfer Learning and M2M-100 Data Augmentation_** The project investigates scalable and effective approaches for improving Semantic Textual Relatedness (STR) in low-resource African languages through transfer learning and multilingual machine translation–based augmentation. ## Key Features - Three African Languages: Hausa, Kinyarwanda, Afrikaans - SemRel2024 Dataset: Standardized benchmark for semantic relatedness - Transfer Learning: Fine-tuning African-centric models (AfriBERTa, AfroXLMR) - M2M-100 Augmentation: MAFAND-MT back-translation pipeline - Significant Gains: Up to **1167% improvement** over baseline models ## 📚 Research Questions **RQ1:** Which transfer-learning methods (AfriBERTa vs AfroXLMR) yield the best STR performance for African languages? **RQ2:** How effective is M2M-100 back-translation in improving model accuracy and robustness under low-resource constraints? ## 📊 Results Summary | Language | Baseline (XLM-R) | Fine-tuned (AfroXLMR) | M2M-100 Augmented | Improvement | |--------------|------------------|-------------------------|-------------------|-------------| | Afrikaans | 0.4016 | 0.4085 | 0.6452 | +60.66% | | Hausa | 0.1247 | 0.6518 | 0.6389 | +412.51% | | Kinyarwanda | 0.0425 | 0.5390 | 0.6456 | +1417.76% | **Metric:** Spearman Correlation Coefficient (ρ) ## 📁 Project Structure ``` ├── COS802_Project_Code.ipynb ├── COS_802_Proposal_Final_u25743695.pdf ├── IEEE_Paper.tex ├── README.md │ ├── results/ │ ├── all_results_mafand.csv │ ├── statistical_tests_mafand.csv │ └── final_comparison_table_mafand.csv │ ├── visualizations/ │ ├── comparison_bar_chart_mafand.html │ ├── improvement …

Visit

github.com

Languages

AfrikaansHausaKinyarwanda

Similaires

clerencemathonsi/semantic-relatedness-african-languagesSemEval Task 1: Semantic Textual Relatedness for African and Asian LanguagesEnhancing African low-resource languages: Swahili data for language modellingSemEval-2024 Task 1: Semantic Textual Relatedness for African and Asian LanguagesGUIDE: Creating Semantic Domain Dictionaries for Low-Resource LanguagesSemRel2024: A Collection of Semantic Textual Relatedness Datasets for 13 Languages

clerencemathonsi/semantic-relatedness-african-languages

# COS802 Semantic Relatedness This repository contains code and resources for experiments on semant

SemEval Task 1: Semantic Textual Relatedness for African and Asian Languages

We present the first shared task on Semantic Textual Relatedness (STR). While earlier shared tasks p

Enhancing African low-resource languages: Swahili data for language modelling

Language modelling using neural networks requires adequate data to guarantee quality word representation which is important for natural language processing (NLP) tasks. However, African languages, Swahili in particular, have been disadvantaged and most of them are

SemEval-2024 Task 1: Semantic Textual Relatedness for African and Asian Languages

This is the GitHub repository hosting data and code for SemEval-2024 Task 1: Semantic Textu

GUIDE: Creating Semantic Domain Dictionaries for Low-Resource Languages

Over 7,000 of the world's 7,168 living languages are still low-resourced. This paper aims to narrow

SemRel2024: A Collection of Semantic Textual Relatedness Datasets for 13 Languages

Exploring and quantifying semantic relatedness is central to representing language and holds signifi