Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Tatoeba – Multilingual Sentence Pairs

Domain:

natural language processing

Record type:

dataset
Machine translation, language modeling, multilingual NLP research Notes / challenges: Contains multiple African languages like Swahili, Yoruba, Zulu, Hausa, Amharic; may require filtering or preprocessing for specific use

Visit

tatoeba.org

Tasks

language modelingmachine translation

Languages

AmharicHausaSwahiliYoruba

Tags

open-data-africaopen-data-africa catalog

Similar

KevinKibe/kikuyu-sentence-pairsYouMike/kikuyu-sentence-pairsUKPLab/conll2020-multilingual-sentence-probingghananlpcommunity/english-ewe-sentence-pairs-4mghananlpcommunity/english-ga-sentence-pairs-400kghananlpcommunity/english-nzema-sentence-pairs-90k

KevinKibe/kikuyu-sentence-pairs

YouMike/kikuyu-sentence-pairs

UKPLab/conll2020-multilingual-sentence-probing

Code and data for our CoNLL 2020 publication: "How to Probe Sentence Embeddings in Low-Resource Lang

ghananlpcommunity/english-ewe-sentence-pairs-4m

This dataset is made available because of Ghana NLP's volunteer driven research work. Please conside

ghananlpcommunity/english-ga-sentence-pairs-400k

This dataset is made available because of Ghana NLP's volunteer driven research work. Please conside

ghananlpcommunity/english-nzema-sentence-pairs-90k

This dataset is made available because of Ghana NLP's volunteer driven research work. Please conside