Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Specialised and general Sepedi corpus

Domain:

natural language processing

Record type:

dataset
Creator:
Mah
Publisher:
University of Pretoria
Host:avatar
The study investigated the syntactic and semantic features of six selected Sepedi conjunctions as observed between Specialised Sepedi corpus and General Sepedi corpus. Furthermore, the study sought to determine whether there are similarities and differences in the usage and meaning of Sepedi conjunctions between scholarly sources and the corpora. The study employed corpus-based approach for data analysis and interpretation, and a corpus-query software called ‘LancsBox X’ was used for querying the corpora. This study was grounded in Noam Chomsky's seminal work on generative grammar. Generative grammar is conceived as a structured system of statements and rules aimed at describing and defining grammatically correct utterances within a language, while excluding those that are not well-formed.

Visit

doi.orgresearchdata.up.ac.za

Languages

Sotho, Northern

Tags

Teacher education and professional development of educators

Licenses

Creative Commons Attribution 4.0 Internationalhttps://creativecommons.org/licenses/by/4.0/legalcode

Similar

Sepedi NER CorpusNCHLT Speech Corpus -- SepediNCHLT Speech Corpus -- SepediLwazi Sepedi TTS corpusLwazi Sepedi ASR corpusNCHLT Sepedi Speech Corpus

Sepedi NER Corpus

The Sepedi Ner Corpus is a Sepedi dataset developed by The Centre for Text Technology (CTexT), North

NCHLT Speech Corpus -- Sepedi

This is the Sepedi language part of the NCHLT Speech Corpus of the South African languages. Language

NCHLT Speech Corpus -- Sepedi

This is the Sepedi language part of the NCHLT Speech Corpus of the South African languages. Language

Lwazi Sepedi TTS corpus

Orthographic and phonemically aligned transcriptions

Lwazi Sepedi ASR corpus

Complete audio recordings and orthographic transcriptions used for Lwazi speech recognition systems.

NCHLT Sepedi Speech Corpus

Orthographically transcribed broadband speech corpus of approximately 56 hours, including a test sui