Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

SFPC (Sango-French Parallel Corpus)

Domain:

natural language processing

Record type:

dataset
Creator:
ala
Host:
The first quality-filtered, verse-aligned Sango-French parallel corpus, constructed for neural machine translation research. This dataset directly addresses the "Sango Problem" identified by Meta's NLLB-200 project — the failure of cross-lingual transfer for a linguistically isolated Creole language. Associated resources: Model: alaminerca/nllb-sango-french Demo: Sango-French Translator

Visit

huggingface.co

Tasks

machine translation

Languages

Sango

Tags

translationparallel-corpusbiblesangoafrican-languageslow-resourcecreolenmt

Licenses

cc-by-4.0

Similar

French-Fongbe Parallel CorpusFrench-Adja Parallel CorpusFrench–Medumba Parallel CorpusEwondo-French Parallel Corpusbesacier/mboshi-french-parallel-corpusBamun-French Parallel Corpus 1.1

French-Fongbe Parallel Corpus

Ce dataset est un corpus parallèle Français-Fongbe (Bénin) généré par IA et structuré pour l'entraîn

French-Adja Parallel Corpus

The first publicly available parallel text corpus for Adja machine translation, targeting an under-r

French–Medumba Parallel Corpus

A small parallel corpus of French ↔ Medumba (byv) sentence pairs, intended as a seed resource for ma

Ewondo-French Parallel Corpus

This dataset is a parallel corpus of Ewondo and French texts. The text was obtained by transcribing

besacier/mboshi-french-parallel-corpus

# mboshi-french-parallel-corpus This repository contains a speech corpus collected during a realist

Bamun-French Parallel Corpus 1.1

This dataset is the second iteration of the 'Bamun–French parallel corpus' that was initially publis