Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Creating a reusable English – Afrikaans parallel corpora for bilingual dictionary construction

Domain:

natural language processing

Record type:

paper
This paper investigates the possibilities in creating a bilingual English – Afrikaans dictionary by building a parallel corpus and using the Uplug tool to process it. The resulting parallel corpus with approximately 400,000 words per language was created partly from texts collected from the South African government and partly from the OPUS corpus. The recall and accuracy of the bilingual dictionary was evaluated based on the statistical data collected. Samples of translations were generated, compiled as questionnaires and then assessed by English – Afrikaans speaking respondents. The results yielded an accuracy of 87.2 percent and a recall of 67.3 percent for the processed dictionary. Our English – Afrikaans parallel corpora can be found at the following address: let.rug.nl

Visit

people.dsv.su.se

Connected records

dataset

Tasks

machine translation

Languages

Afrikaans

Licenses

Similar

Autshumato English-Afrikaans Parallel CorporaUsing ParaConc to extract bilingual terminology from parallel corpora: A case of English and NdebeleAutshumato English-Xitsonga Parallel CorporaAutshumato English-Setswana Parallel CorporaAutshumato English-Sepedi Parallel CorporaAutshumato English-isiZulu Parallel Corpora

Autshumato English-Afrikaans Parallel Corpora

Aligned parallel corpora for the language pair English-Afrikaans. The data is given as two separate

Using ParaConc to extract bilingual terminology from parallel corpora: A case of English and Ndebele

The development of African languages into languages of science and technology is dependent on action

Autshumato English-Xitsonga Parallel Corpora

Aligned English-Xitsonga parallel corpus. The data is given as two seperate UTF-8 text files; with e

Autshumato English-Setswana Parallel Corpora

Aligned English-Setswana parallel corpus. This set contains data that was translated by professional

Autshumato English-Sepedi Parallel Corpora

Aligned parallel corpora for the language pair English-Sepedi. The data is given as two separate UTF

Autshumato English-isiZulu Parallel Corpora

Aligned parallel corpora for the language pair English-isiZulu. The data is given as two separate UT