Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Back-translated Swahili-French 1M sentence parallel data

Domaine:

natural language processing

Type de record:

dataset
Synthetic data used in the experiments of the paper "Congolese Swahili Machine Translation for Humanitarian Response" published in Africa NLP workshop organized within the 16th Conference of the European Chapter of the Association for Computational Linguistics (EACL2021)

Visit

gamayun.translatorswb.org

Connected records

paper

Tasks

machine translation

Languages

Swahili

Tags

clear global

Similaires

TWB Parallel Sentence kits - Swahili (5k)TWB Parallel Sentence kits - Congo Swahili (25k)Fon French Daily Dialogues Parallel DataSMOL: Professionally translated parallel data for 115 under-represented languagesEnglish-Giriama Parallel Sentence DatasetNekonardo/lrl-parallel-sentence-mining

TWB Parallel Sentence kits - Swahili (5k)

The Swahili portion of CLEAR Global's Gamayun Language Data Kits — 5,000 parallel English–Swahili se

TWB Parallel Sentence kits - Congo Swahili (25k)

The Congo Swahili portion of CLEAR Global's Gamayun Language Data Kits — 25,305 parallel French–Cong

Fon French Daily Dialogues Parallel Data

We aim to collect, clean, and store corpora of Fon and French sentences for Natural Languag

SMOL: Professionally translated parallel data for 115 under-represented languages

We open-source SMOL (Set of Maximal Overall Leverage), a suite of training data to unlock machine tr

English-Giriama Parallel Sentence Dataset

This dataset consists of sentence pairs in English and their corresponding translations in Giriama (

Nekonardo/lrl-parallel-sentence-mining

Codebase for benchmarking and enhancing multilingual sentence embeddings for parallel sentence minin