A Zarma-French parallel corpus for Machine Translation
# Feriji: A French-Zarma Parallel Corpus, Glossary & Translator
This repository contains Feriji, a work-in-progress French-Zarma parallel corpus curated by **Habibatou Abdoulaye Alfari**, **Elysabhete Amadou Ibrahim**, **Christopher Homan**, and **Mamadou K. KEITA**. Feriji is a collection of 61,085 aligned machine translation-ready French-Zarma lines curated from various sources. The corpus aims to contribute to the development of machine translation systems and linguistic studies between French and Zarma languages.
## Dataset Description
- **Size**: 61,085 sentences in Zarma and 42,789 in French.
- **Glossary**: 4,062 words.
## Dataset Statistics
| | French | Zarma |
|-------------------|-----------|----------|
| Number of sentences | 42,789 | 61,085 |
| Glossary entries | 4,062 | 4,062 |
| Unique words | 21,592 | 9,902 |
## Usage
The dataset is intended for academic research and development of Machine Translation systems. You can test the Feriji Translator here.
## Acknowledgements
We would like to thank our institutions, especially Ashesi University, and contributors for their support in creating this resource. The Computer Science department of Ashesi University provided financial and cloud resources support, which was crucial for this project.
## Citations
If you use this dataset in your research, please cite it as follows:
```bibtex
@dataset{Feriji,
author = {Habibatou Abdoulaye Alfari, Elysabhete Amadou Ibrahim, Christopher Homan, and Mamadou K. KEITA},
title = {Feriji: A French-Zarma Parallel Corpus, Glossary & Translator},
year = 2023,
publisher = {GitHub},
journal = {GitHub repository},
howpublished = {\url{
github.com
}