Logo Lanfrica

27-GROUP/Feriji

Domaine:

natural language processing

Type de record:

dataset
Créateur:
27-
Hôte:
A Zarma-French parallel corpus for Machine Translation # Feriji: A French-Zarma Parallel Corpus, Glossary & Translator This repository contains Feriji, a work-in-progress French-Zarma parallel corpus curated by **Habibatou Abdoulaye Alfari**, **Elysabhete Amadou Ibrahim**, **Christopher Homan**, and **Mamadou K. KEITA**. Feriji is a collection of 61,085 aligned machine translation-ready French-Zarma lines curated from various sources. The corpus aims to contribute to the development of machine translation systems and linguistic studies between French and Zarma languages. ## Dataset Description - **Size**: 61,085 sentences in Zarma and 42,789 in French. - **Glossary**: 4,062 words. ## Dataset Statistics | | French | Zarma | |-------------------|-----------|----------| | Number of sentences | 42,789 | 61,085 | | Glossary entries | 4,062 | 4,062 | | Unique words | 21,592 | 9,902 | ## Usage The dataset is intended for academic research and development of Machine Translation systems. You can test the Feriji Translator here. ## Acknowledgements We would like to thank our institutions, especially Ashesi University, and contributors for their support in creating this resource. The Computer Science department of Ashesi University provided financial and cloud resources support, which was crucial for this project. ## Citations If you use this dataset in your research, please cite it as follows: ```bibtex @dataset{Feriji, author = {Habibatou Abdoulaye Alfari, Elysabhete Amadou Ibrahim, Christopher Homan, and Mamadou K. KEITA}, title = {Feriji: A French-Zarma Parallel Corpus, Glossary & Translator}, year = 2023, publisher = {GitHub}, journal = {GitHub repository}, howpublished = {\url{github.com }