Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Low resource Twi-English parallel corpus for machine translation in multiple domains (Twi-2-ENG)

Domain:

natural language processing

Record type:

dataset
Creator:
EmmXiaSteAma
Publisher:
Spr
Host:
Abstract Although Ghana does not have one unique language for its citizens, the Twi dialect stands a chance of fulfilling this purpose. Twi is among the low-resourced language categories, yet it is widely spoken beyond Ghana and in countries such as the Ivory Coast, Benin, Nigeria, and other places. However, it continues to be seen as the perfect resource for Twi Machine Translation (MT) of IS0 639-3. The issue with the Twi-English parallel corpus is eminent at the multiple domain dataset level, partly due to the complex design structure and scarcity of the digital Twi lexicon. This study introduced Twi-2-ENG, a large-scale multiple domain Twi to English parallel corpus, Twi digital Dictionary, and lexicon version of Twi. Also, it employed the Ghanaian Parliamentary Hansards, crowdsourcing, and digital Ghana News Portals to crawl all the English sentences. Our curled news portals accumulated 5,765 parallel corpus sentences, the Twi New Testament Bible, and social media platforms. The data-gathering method used means of translation, compilation, tokenization, and the final alignments with the Twi-English parallel sentences, including the technology employed in compiling and hosting the corpus, were duly discussed. The results reveal that the role of manually qualified linguistic professionals and Twi translation specialists across the media spectrum, academia, and well-wishers adds a considerable volume to the Twi-2-ENG parallel corpus. Finally, all the sentences were curated with the help of a corpus manager, sketch engine, linguistics, and professional translators to align and tokenize all texts, allowing the Twi professional linguists to evaluate the corpus.

Visit

doi.org

Tasks

machine translation

Languages

AkanBwamu, CwiDinka, SoutheasternTwi

Licenses

https://creativecommons.org/licenses/by/4.0https://creativecommons.org/licenses/by/4.0

Similar

English-Twi Parallel Corpus for Machine TranslationTWIENG: A Multi-Domain Twi-English Parallel Corpus for Machine Translation of Twi, a Low-Resource African LanguageTwieng: A Multi-Domain Twi-English Parallel Corpus for Machine Translation of the Twi Language, A Low-Resource African LanguageTwi–English Parallel Corpus for AgricultureENGLISH-AKUAPEM TWI PARALLEL CORPUSTwi-English Citizens' Budget Parallel Corpus

English-Twi Parallel Corpus for Machine Translation

We present a parallel machine translation training corpus for English and Akuapem Twi of 25,421 sentence pairs. We used a transformer-based translator to generate initial translations in Akuapem Twi, which were later verified and corrected where necessary by native

TWIENG: A Multi-Domain Twi-English Parallel Corpus for Machine Translation of Twi, a Low-Resource African Language

A Twi-English parallel corpus is certainly an important resource for Machine Translation of Twi (ISO

Twieng: A Multi-Domain Twi-English Parallel Corpus for Machine Translation of the Twi Language, A Low-Resource African Language

A Twi-English parallel corpus is certainly an important resource for Machine Translation of Twi (ISO

Twi–English Parallel Corpus for Agriculture

This is a parallel corpus of Twi (Akan) transcriptions and their English translations, focused on th

ENGLISH-AKUAPEM TWI PARALLEL CORPUS

This dataset (verified_data.csv) is bilingual machine translation training corpus for English and Akuapem Twi of 25,421 sentence pairs. 
A transformer-based machine translator was u

Twi-English Citizens' Budget Parallel Corpus

317 Asante Twi ↔ English parallel sentences in the economy / public-finance domain, extracted and al