Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Multilingual Data from the Agricultural Domain: Presenting the NWU-Pula/Imvula Corpora

Domain:

natural language processingagriculture

Record type:

dataset
Creator:
TanCindy McKellarMar
Publisher:
Dig
Host:
This paper presents new multilingual corpora from the agricultural domain for seven South African Languages, namely Afrikaans, English, isiXhosa, isiZulu, Sesotho, Sesotho sa Leboa, and Setswana, based on the Pula/Imvula magazine. After pre-processing, the data has been automatically sentencized, tokenized, lemmatized and annotated with part-of-speech information using the services available at v-ctx-lnx7.nwu.ac.za. The final resources comprising between 774k and 1,38M tokens per language are included on the Corpus Cooperative at North-West University (COCO@NWU) corpus platform at coco.nwu.ac.za as searchable corpora. In addition, the data can be made avail- able as text files for research purposes upon request. To highlight the value of this agricultural domain-specific data collection in relation to more general data, we also include some corpus-based statistics and comparisons with previous research.

Visit

doi.org

Languages

AfrikaansSetswanaSotho, NorthernSotho, SouthernXhosaZulu

Licenses

https://creativecommons.org/licenses/by-sa/4.0

Similar

Development of Multilingual Corpora in Medical Domain Using Neural Machine TranslationSmall-Multilingual-CorporaPreparing the Vuk'uzenzele and ZA-gov-multilingual South African multilingual corporamultilingual Corpora for Ethiopian LanguagesEllienware/PulaOxxoCodes/Pula

Development of Multilingual Corpora in Medical Domain Using Neural Machine Translation

Development of Multilingual Corpora in Medical Domain Using Neural Machine Translation

Poster presented at the Deep Learning Indaba 2022 by AJAGBE, Sunday Adeola

Small-Multilingual-Corpora

Small Multilingual Pretraining Copora used in the ICML 2025 Paper: Banyan: Improved Representation L

Preparing the Vuk'uzenzele and ZA-gov-multilingual South African multilingual corpora

This paper introduces two multilingual government themed corpora in various South African languages.

multilingual Corpora for Ethiopian Languages

Ellienware/Pula

Pula is an AI-powered financial and mobility ecosystem built for underserved communities in South Af

OxxoCodes/Pula

Pula: The first suite of LLMs for Setswana