Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Corpus Construct: A Research Tool in Syntactic Analysis of Bantu Languages

Domain:

natural language processing

Record type:

dataset
Creator:
Wal
Publisher:
Uta
Host:
Syntax of languages is understood to be shaped by syntactic universals principles. Despite operating within the constraints of these universals, languages have managed to display their unique syntactic features. This is explained by specificity in the ranking of these universals across languages. To establish the unique features, researchers have to use data collected using a variety of methods some of which are corpus studies, linguistic elicitation, introspection and experimentation. Each of these methods requires a research tool whose development or adoption is dependent on the research question(s)formulated to fill a study gap. The use of corpus construct to generate data, as other research tools, requires an understanding of what type of data is needed to answer which questions on which linguistic features. The construction of a corpus can be done from plain texts or annotated texts. The question is: how can corpus construct be used as a tool in the study of languages whose corpus is yet to be compiled and made available online? This article, therefore, intends to answer this question with biases on Bantu languages. It will be necessary to make databases from the corpus constructs available by building corpora for the languages in question. The findings of this article are deemed important in offering knowledge on building of corpus and how to use the built corpus to investigate a syntactic feature in a Bantu language.

Visit

doi.org

Tasks

parsing

Similar

Towards a unified theory of morphological productivity in the Bantu languages: A corpus analysis of nominalization patterns in SwahiliA corpus-based analysis of P indexing in Ruuli (Bantu, JE103)A Comparative Discourse Analysis of Some Bantu Languages in TanzaniaExploring the Impact of Academic Writing Expertise on Syntactic Complexity: a Corpus-Based AnalysisReview article: Second language acquisition of Bantu languages: A (mostly) untapped research opportunityLexical cluster analysis of 10 Bantu A80 languages

Towards a unified theory of morphological productivity in the Bantu languages: A corpus analysis of nominalization patterns in Swahili

Models arguing for a connection between morphological productivity and relative morpheme frequency have focused on languages with relatively low average morpheme to word ratios. Typologically synthetic languages like Swahili which have relatively high average morph

A corpus-based analysis of P indexing in Ruuli (Bantu, JE103)

A Comparative Discourse Analysis of Some Bantu Languages in Tanzania

https://www.sil.org/resources/archives/57053

Exploring the Impact of Academic Writing Expertise on Syntactic Complexity: a Corpus-Based Analysis

Abstract: Intense debates in foreign or second language learning evolve around English language acad

Review article: Second language acquisition of Bantu languages: A (mostly) untapped research opportunity

This review article presents a summary of research on the second language acquisition of Bantu langu

Lexical cluster analysis of 10 Bantu A80 languages

International audience In this presentation, I will show the results of a small-scale