Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

The SAWA Corpus: a Parallel Corpus English - Swahili

Domaine:

natural language processing

Type de record:

paper
Research in data-driven methods for Machine Translation has greatly benefited from the increasing availability of parallel corpora. Processing the same text in two different languages yields useful information on how words and phrases are translated from a source language into a target language. To investigate this, a parallel corpus is typically aligned by linking linguistic tokens in the source language to the corresponding units in the target language. An aligned parallel corpus therefore facilitates the automatic development of a machine translation system and can also bootstrap annotation through projection. In this paper, we describe data collection and annotation efforts and preliminary experimental results with a parallel corpus English - Swahili.

Visit

aclanthology.org

Tasks

machine translation

Languages

Swahili

Tags

sawa

Similaires

Akan–English Maternal Health Parallel Corpus Akan–English Maternal Health Parallel CorpusSomali-English Parallel CorpusNupe-English parallel corpusAmharic-English Parallel Corpusluganda-english-parallel-corpusEnglish Hausa Parallel Corpus

Akan–English Maternal Health Parallel Corpus Akan–English Maternal Health Parallel Corpus

This dataset contains a curated bilingual parallel corpus developed to support domain-specific neura

Somali-English Parallel Corpus

This dataset contains high-quality parallel sentence pairs, multi-sentence alignments, and paragraph

Nupe-English parallel corpus

This is the first ever Nupe - English Parallel Corpus and Nupe Monolingual Corpora curated from diverse sources including poems,idioms, proverbs, religpoius text etc. The aim of this data collection is to make available a cultural-aware Nupe-english corpus for NLP

Amharic-English Parallel Corpus

This corpus consists of 145,820 Amharic-English parallel sentences (segments) from various sources. This corpus is larger in size than previously compiled corpora. It is released for research purposes and can be used to train or support Amharic-English machine tran

luganda-english-parallel-corpus

English Hausa Parallel Corpus

This English–Hausa Parallel Corpus is a curated bilingual dataset of 5,000 aligned sentence pairs, t