Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

TaPaCo: A Corpus of Sentential Paraphrases for 73 Languages

Domaine:

natural language processing

Type de record:

datasetpaper
A freely available paraphrase corpus for 73 languages extracted from the Tatoeba database. Tatoeba is a crowdsourcing project mainly geared towards language learners. Its aim is to provide example sentences and translations for particular linguistic constructions and words. The paraphrase corpus is created by populating a graph with Tatoeba sentences and equivalence links between sentences “meaning the same thing”. This graph is then traversed to extract sets of paraphrases. Several language-independent filters and pruning steps are applied to remove uninteresting sentences. A manual evaluation performed on three languages shows that between half and three quarters of inferred paraphrases are correct and that most remaining ones are either correct but trivial, or near-paraphrases that neutralize a morphological distinction. The corpus contains a total of 1.9 million sentences, with 200 – 250 000 sentences per language. It covers a range of languages for which, to our knowledge, no other paraphrase dataset exists.

Visit

zenodo.org

Tasks

natural language generationmachine translationembeddings

Languages

AfrikaansAmazighRundi

Licenses

Creative Commons Attribution 2.0 Generic

Similaires

A Corpus for Berber LanguagesCross-Lingual Query Generation Augmentation for Robust Dense Retrieval Against Adversarial Paraphrases in Low-Resource LanguagesDeveloping a Model Corpus for Endangered LanguagesAfroMAFT Corpus: Language Adaptation Corpus for African languagesZambezi Voice: A Multilingual Speech Corpus for Zambian Languagesemmanuel2406/web-corpus-for-african-languages

A Corpus for Berber Languages

International audience

Cross-Lingual Query Generation Augmentation for Robust Dense Retrieval Against Adversarial Paraphrases in Low-Resource Languages

Effective cross-lingual dense retrieval methods that rely on multilingual pre-trained language model

Developing a Model Corpus for Endangered Languages

World languages are becoming endangered at an unprecedented rate. Linguists estimate that 50-90% of

AfroMAFT Corpus: Language Adaptation Corpus for African languages

Language Adaptation Corpus for 17 African languages, English, French, and Arabic.

Zambezi Voice: A Multilingual Speech Corpus for Zambian Languages

This work introduces Zambezi Voice, an open-source multilingual speech resource for Zambian languages. It contains two collections of datasets: unlabelled audio recordings of radio news and talk shows programs (160 hours) and labelled data (over 80 hours) consistin

emmanuel2406/web-corpus-for-african-languages

# Web Corpus for african languages Created by Emmanuel Rassou Info Doc can be accessed here **War