Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Using ParaConc to extract bilingual terminology from parallel corpora: A case of English and Ndebele

Domaine:

natural language processing

Type de record:

datasetpaper
Créateur:
Ket
Éditeur:
AOS
Hôte:
The development of African languages into languages of science and technology is dependent on action being taken to promote the use of these languages in specialised fields such as technology, commerce, administration, media, law, science and education among others. One possible way of developing African languages is the compilation of specialised dictionaries (Chabata 2013). This article explores how parallel corpora can be interrogated using a bilingual concordancer (ParaConc) to extract bilingual terminology that can be used to create specialised bilingual dictionaries. An English–Ndebele Parallel Corpus was used as a resource and through ParaConc, an alphabetic list was compiled from which headwords and possible translations were sought. These translations provided possible terms for entry in a bilingual dictionary. The frequency feature and ‘hot words’ tool in ParaConc were used to determine the suitability of terms for inclusion in the dictionary and for identifying possible synonyms, respectively. Since parallel corpora are aligned and data are presented in context (Key Word in Context), it was possible to draw examples showing how headwords are used. Using this approach produced results quickly and accurately, whilst minimising the process of translating terms manually. It was noted that the quality of the dictionary is dependent on the quality of the corpus, hence the need for creating a representative and clean corpus needs to be emphasised. Although technology has multiple benefits in dictionary making, the research underscores the importance of collaboration between lexicographers, translators, subject experts and target communities so that representative dictionaries are created.

Visit

doi.org

Languages

NdebeleNdebele

Licenses

https://creativecommons.org/licenses/by/4.0

Similaires

Creating a reusable English – Afrikaans parallel corpora for bilingual dictionary constructionAnalysing the English-Xhosa parallel corpus of technical texts with Paraconc: a case study of term formation processesAutshumato English-Xitsonga Parallel CorporaAutshumato English-Setswana Parallel CorporaAutshumato English-Afrikaans Parallel CorporaAutshumato English-Sepedi Parallel Corpora

Creating a reusable English – Afrikaans parallel corpora for bilingual dictionary construction

This paper investigates the possibilities in creating a bilingual English – Afrikaans dictionary by building a parallel corpus and using the Uplug tool to process it. The resulting parallel corpus with approximately 400,000 words per language was created partly fro

Analysing the English-Xhosa parallel corpus of technical texts with Paraconc: a case study of term formation processes

Autshumato English-Xitsonga Parallel Corpora

Aligned English-Xitsonga parallel corpus. The data is given as two seperate UTF-8 text files; with e

Autshumato English-Setswana Parallel Corpora

Aligned English-Setswana parallel corpus. This set contains data that was translated by professional

Autshumato English-Afrikaans Parallel Corpora

Aligned parallel corpora for the language pair English-Afrikaans. The data is given as two separate

Autshumato English-Sepedi Parallel Corpora

Aligned parallel corpora for the language pair English-Sepedi. The data is given as two separate UTF