Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

mjameeel/finetuning_mT5_for_english_to_hausa_translation

Domaine:

natural language processing

Type de record:

softwareproject
Créateur:
mja
Hôte:
Fine-tuning Google’s mT5 model for English–Hausa machine translation, with an end-to-end pipeline covering data preparation, and model training. # Fine-Tuning mT5 for English–Hausa Translation This repository contains an implementation of **fine-tuning the multilingual T5 (mT5) model for English to Hausa machine translation**, along with an inference. The project demonstrates how to prepare a bilingual dataset, fine-tune a pretrained sequence-to-sequence transformer using the Hugging Face ecosystem, and deploy the resulting model for real-world usage. --- ## 📌 Project Overview - **Task:** Neural Machine Translation (English → Hausa) - **Model:** `google/mt5-small` - **Frameworks:** Hugging Face Transformers, Datasets, PyTorch - **Execution Environment:** Google Colab (Free GPU) This project is particularly relevant for **low-resource language translation** and can be extended to other African or multilingual NLP tasks. --- ## 🧠 Key Features - End-to-end fine-tuning of an mT5 model - Dataset preparation and sampling - Tokenization for sequence-to-sequence learning - Model training and inference - Translation of unseen English sentences --- ## 📂 Repository Structure ├── Finetunining_mT5.ipynb # Main notebook (training + inference) ├── requirements.txt # Project dependencies ├── translation.jpg # Project illustration └── README.md # Project documentation --- ## 📊 Dataset - **Source:** English–Hausa parallel corpus from Kaggle - **Format:** CSV (`en-ha.csv`) - **Preprocessing Steps:** - Removal of unnecessary columns - Random sampling for efficient training - Conversion to Hugging Face `Dataset` object The dataset is loaded programmatically using `kagglehub` and prepared for transformer-based training. --- ## ⚙️ Environment Setup The project is designed to run seamlessly on **Google Colab** with GPU acceleration. ### Install Dependencies pip install -U transformers datasets sentencepiece accelerate kagglehub --- ## 🚀 Future Improvements - Train on larger and more diverse datasets - Add BLEU / chrF evaluation metrics - Deploy using Hugging Face Spaces or Docker - Extend to bidirectio …

Visit

github.com

Tasks

machine translation

Languages

Hausa