Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Coptic–French NMT Dataset

Domaine:

natural language processing

Type de record:

dataset
Créateur:
cha
Hôte:
This dataset was created and used as part of the research article: TBA The Coptic source text is derived from digitized editions of Biblical texts. The French side includes multiple translations (e.g., Louis Segond, Darby, Crampon), depending on the experiment. All texts have been normalized and filtered for length and content. See the original paper for full preprocessing and data preparation details. 📊 Usage

Visit

huggingface.co

Tasks

machine translation

Languages

Coptic

Similaires

Adeptschneider/dyula-to-french-nmt-tensorflowDataset Curation for Kalabari NMT SystemFon-French DatasetWolof-French ASR DatasetBambara French Parallel datasetA Text Recognition Dataset from Sahidic Coptic Ancient Manuscripts

Adeptschneider/dyula-to-french-nmt-tensorflow

Dataset Curation for Kalabari NMT System

The development of Natural Language Processing (NLP) tools for endangered and low resource languages

Fon-French Dataset

FFR Dataset is an ongoing project to collect, clean and store corpora of Fon and French sentences for machine translation from Fon-French. Fon (also called Fongbe) is an African-indigenous language spoken mostly in Benin, by about 1.7 million people. As training da

Wolof-French ASR Dataset

Dataset unifié pour l'entraînement de modèles de reconnaissance automatique de la parole (ASR) en wo

Bambara French Parallel dataset

Paired Bambara and French texts for Natural Language Processing

A Text Recognition Dataset from Sahidic Coptic Ancient Manuscripts

In this work, we target Handwritten Text Recognition (HTR) in low-resource scenarios, which arise fr