Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Bamun-French Parallel Corpus 2.0

Domain:

natural language processing

Record type:

dataset
Creator:
Ins
Host:
This dataset is an extended and updated version of the 'Bamun-French Parallel Corpus 1.1' that is published on the Mozilla Data Collective platform. It is a parallel corpus of 4,444 lines in Bamun and French suitable for machine translation tasks. The text was obtained by transcribing raw audio files. Translations were added to enrich the original corpus. Bamun and French text alignment was performed in the process of creating this dataset. This version of the dataset resolves formatting issues flagged in the original and nearly doubles the number of aligned translation units compared to version 1.1.

Visit

mozilladatacollective.com

Tasks

machine translation

Languages

Bamun

Tags

mdcmozilla data collectiveMTTSV

Licenses

Nwulite Obodo Open Data Licence 1.0 (NOODL-1.0)

Similar

Bamun-French Parallel Corpus 1.1French–Medumba Parallel CorpusEwondo-French Parallel CorpusFrench-Fongbe Parallel CorpusFrench-Adja Parallel CorpusMada-French Parallel Corpus 1.0

Bamun-French Parallel Corpus 1.1

This dataset is the second iteration of the 'Bamun–French parallel corpus' that was initially publis

French–Medumba Parallel Corpus

A small parallel corpus of French ↔ Medumba (byv) sentence pairs, intended as a seed resource for ma

Ewondo-French Parallel Corpus

This dataset is a parallel corpus of Ewondo and French texts. The text was obtained by transcribing

French-Fongbe Parallel Corpus

Ce dataset est un corpus parallèle Français-Fongbe (Bénin) généré par IA et structuré pour l'entraîn

French-Adja Parallel Corpus

The first publicly available parallel text corpus for Adja machine translation, targeting an under-r

Mada-French Parallel Corpus 1.0

This dataset comprises a parallel corpus of Mada–French literary text translations totalling 2,154 l