# Ngambay-French Neural Machine Translation
Welcome to the Ngambay-French Neural Machine Translation project repository. Our project marks a significant step towards bridging language barriers by enabling translation from Ngambay, a Central African language, to French, one of the world's widely spoken languages. We've harnessed the power of state-of-the-art pretrained models like M2M100, T5, and ByT5 to achieve this. You can find the paper detailing our work on ArXiv by Sakayo Toadoum Sari and Angela Fan and Lema Logamou Seknewna.
## The Masakhane Initiative
Our project is part of the broader Masakhane initiative (Masakhane.io), an inspiring online community of African researchers dedicated to advancing machine translation for African languages. While Masakhane has contributed numerous translation models and baselines for various African languages, "Ngambay" language have not yet been explored before.
## Data Sourcing and Compilation
Sourcing the data for this research was a meticulous process, crucial to the success of our project. Our researchers employed a combination of "web-scraping" and "parsing" techniques to gather data from open-source dataset websites. This diligent effort resulted in the collection of over 31,000 Ngambay-French parallel words and sentences, forming the foundation of our project's pilot stage.
## Data Cleaning and Preprocessing
To ensure the accuracy and effectiveness of our Ngambay-French translation model, we subjected the dataset to rigorous cleaning, preprocessing, and tokenization. Special attention was given to preserving the diacritics and special characters unique to the Ngambay alphabet. This meticulous approach enhances the quality of the translations and ensures that the neural model can handle the intricacies of the Ngambay language effectively.
We're excited about the potential impact of this project and look forward to furthering research in machine translation for African languages. Feel free to explore our paper for …