Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

travisfoundation/Tigrinya-Parallel-Corpus

Domaine:

natural language processing

Type de record:

dataset
Créateur:
tra
Hôte:
A Collection of Parallel English-Tigrinya Translations # The TravisFoundation English-Tigrinya Parallel Corpus A corpus of parallel English-Tigrinya phrases, compiled by Travis Foundation. This parallel corpus' main goal is of course to help progress the path towards English-Tigrinya machine translation. Its secondary aim is to increase (if only by a little) the amount of digital, publicly available samples of the Tigrinya language, spoken by at least 9 million people. ## Basic Properties This parallel corpus uses the simplest possible format, namely a line-by-line format in which line `i` in `some_file_EN` corresponds to line `i` in `some_file_TI`. All files are UTF-8 encoded (Unix line-breaking), which includes the Ge'ez alphabet, the alphabet in which Tigrinya is written. The corpus is left unchanged from its original sources (see section below) which entails partly inconsistent style, e.g. with respect to punctuation, and some linguistic noise. This repository includes simple scripts to normalise the corpus linguistically and regarding punctuation, which may be less useful for the English parts of this corpus but more so for the Tigrinya parts. Some cornerstone statistics of the Travis Foundation corpus: - number of phrases: - number of words: EN, TI - | | TI | EN | |-------| ------------- | ------------- | | number of words | 10k | 20k | | number of characters | 100k | 200k | ## Provenance The majority of this collection of parallel sentences is actually scraped from the Jehova's Witnesses' website, to our knowledge the largest findable collection of texts with both English and Tigrinya version. The X parallel sentences obtained from this source account for Y% of our corpus. We have included this data in our corpus simply for its size but warn against two aspects: First, the data may have some problems in linguistic terms. As an English-Tigrinya parallel corpus, we do not know the process by which the original English texts were translated into Tigrinya and we therefore have no grip on the close …

Visit

github.com

Tasks

machine translation

Languages

Tigrigna

Licenses

CC-BY-SA-4.0

Similaires

parallel corpusTWB Parallel Sentence kits - Tigrinya (5k)msquarme/Parallel-CorpusSIMBA9657/haddas-tigrinya-corpusAkan–English Maternal Health Parallel Corpus Akan–English Maternal Health Parallel CorpusEnglish Hausa Parallel Corpus

parallel corpus

A 48,000 Kinyarwanda English Parallel datasets for machine translation, made by curating and transla

TWB Parallel Sentence kits - Tigrinya (5k)

The Tigrinya portion of CLEAR Global's Gamayun Language Data Kits — 5,000 parallel English–Tigrinya

msquarme/Parallel-Corpus

Scraping JW.org to create a parallel corpus for Tigrigna - English and Amharic - English # Parall

SIMBA9657/haddas-tigrinya-corpus

Monolingual Tigrinya newspaper text segmented into article bodies. Suitable for continued pretrainin

Akan–English Maternal Health Parallel Corpus Akan–English Maternal Health Parallel Corpus

This dataset contains a curated bilingual parallel corpus developed to support domain-specific neura

English Hausa Parallel Corpus

This English–Hausa Parallel Corpus is a curated bilingual dataset of 5,000 aligned sentence pairs, t